The work
You receive model outputs on non-trivial security problems: a root-cause analysis of a memory corruption bug, a proposed exploit chain against a hardened target, a triage verdict on a vulnerability report, a patch recommendation. Your job is to decide whether the reasoning actually holds — whether the primitive described is genuinely reachable, whether mitigations like ASLR, CFI, or sandboxing were accounted for, whether the "fix" closes the bug class or just the reported path. Plausible-sounding but wrong is the dominant failure mode, and catching it is the entire value of the role.
Beyond grading, contributors help build benchmark tasks: constructing scenarios with unambiguous ground truth, writing rubrics that distinguish a working exploit from a narrative about one, and flagging where the current evaluation scheme rewards confident prose over verified execution. Domains rotate across binary exploitation, Linux and Windows internals, browser security, web application security, cloud and container misconfiguration, cryptographic implementation flaws, and secure code review.
What the screen measures
Mercor's AI interview probes credentials for specificity, then follows up on depth. Expect to be asked which CVE you found, how you found it, what the primitive was, and what the vendor pushed back on — vague answers collapse quickly under a second question. Public artifacts matter here more than most categories: CVE IDs, CTF team and event placements, hall-of-fame listings, conference papers, or maintained tooling. Writing quality is also assessed directly, since the deliverable is written justification another researcher can audit.
Logistics
- Fully remote and asynchronous; work is claimed from a queue rather than scheduled.
- Most contributors commit 10–20 hours per week, with variable volume between project phases.
- Observed rates for this listing run $200–250/hr, set per project and not guaranteed; final rate depends on screening outcome and task tier.
- Expect a paid or unpaid calibration exercise before full task access, and periodic rubric updates as benchmarks evolve.