Build one Dockerized Python scenario to evaluate safe agent code fixes
Project brief
Looking for a benchmark author, not a prompt writer or annotator. Build one self-contained coding scenario that tests whether an AI agent fixes a security problem properly rather than making a cosmetic patch.
The package needs Python repository fixtures, clear task instructions, a working reference fix and a deterministic verifier, all runnable in Docker. Review a test agent trajectory and explain any failure you find.
Strong Python, Docker and security-fix experience matter. The scenario and authorized or synthetic code will be agreed before starting; keep testing isolated and local.
Deliverables & acceptance
What you'll deliver
- Runnable Docker scenario with fixtures, instructions, reference fix and Python verifier
- Brief trajectory review and reproducible run instructions
What the result must meet
- A clean container run passes the reference fix and rejects the unfixed fixture and an agreed superficial patch
- Documented commands reproduce verifier results; trajectory review identifies the failure and its cause