What the work looks like
You receive developer traces: a task, a repo state, and the complete record of an AI-assisted session — prompt turns, file diffs, terminal output, test runs, and the developer's or agent's reasoning along the way. Your job is to decide whether the end state is actually correct (not just green tests), whether the path taken was a reasonable engineering path, and where the trajectory went wrong if it did. That means reading unfamiliar code fast, reproducing or mentally executing changes, spotting silent regressions, and catching the classic agentic failure modes: tests weakened to pass, error swallowed in a try/except, a stub left in place, scope creep across twenty files when three would do.
Output is written. Each audit pairs a rubric score with prose that a model trainer can act on — cite the specific turn, the specific line, and what a competent engineer would have done instead. Vague verdicts ("code quality is poor") are the most common reason submitted work gets rejected. Expect calibration rounds against gold-standard traces, occasional disagreement resolution with other auditors, and rubric updates you're expected to absorb mid-project.
What the screen looks for
- Verifiable professional development history — three years minimum, with languages and stacks you can be interrogated about.
- Real fluency with agentic and spec-driven tooling: Cursor, Claude Code, Copilot agent mode, Kiro, or similar. Not "I've tried it" — what you shipped with it and where it failed you.
- Code-reading under time pressure across full-stack or backend systems you didn't write.
- Judgment: can you separate a stylistic preference from a correctness bug, and hold that line under follow-up questioning?
Logistics
Fully remote and asynchronous, contract via Mercor. Observed rates for this band sit at $70–90/hr, set by experience and calibration performance — not guaranteed, and typically tied to a weekly hour commitment agreed at onboarding (often 10–20 hours, sometimes full-time on active projects). Work arrives in batches with turnaround windows measured in days rather than hours, so you set your own schedule, but responsiveness to rubric clarifications during a live project matters. Audits are throughput-tracked and quality-sampled; sustained low agreement with gold labels ends engagements.