← All tasks
Custom software development for businessPay for results

Build one Dockerized Python scenario to evaluate safe agent code fixes

Project brief

Looking for a benchmark author, not a prompt writer or annotator. Build one self-contained coding scenario that tests whether an AI agent fixes a security problem properly rather than making a cosmetic patch.

The package needs Python repository fixtures, clear task instructions, a working reference fix and a deterministic verifier, all runnable in Docker. Review a test agent trajectory and explain any failure you find.

Strong Python, Docker and security-fix experience matter. The scenario and authorized or synthetic code will be agreed before starting; keep testing isolated and local.

Deliverables & acceptance

What you'll deliver

  • Runnable Docker scenario with fixtures, instructions, reference fix and Python verifier
  • Brief trajectory review and reproducible run instructions

What the result must meet

  • A clean container run passes the reference fix and rejects the unfixed fixture and an agreed superficial patch
  • Documented commands reproduce verifier results; trajectory review identifies the failure and its cause
Browse more tasks →How pay-for-results hiring works →

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.