What the work involves
You are the legal substance behind a frontier model's training pipeline. Day to day, that means reviewing legal knowledge-work tasks and model outputs for the failure modes only a practitioner catches — an argument that cites the right doctrine but misses the controlling standard, a memo that reads persuasively and would still get you called into a partner's office. You will write instruction specs that define what "correct" looks like for a given legal task, produce golden solutions yourself, and design benchmark sets hard enough to expose real weakness rather than confirm progress.
The calibration side matters as much as the drafting. You will work alongside client researchers and experts in adjacent domains to convert tacit legal judgment — the reason a senior associate rejects a first draft — into explicit, teachable criteria that other reviewers can apply consistently. Expect substantial written output and close collaboration inside the client's own tooling, on client-issued accounts and equipment.
What the screen looks for
Mercor's screening is AI-led and follow-up heavy. It is built to distinguish a practising specialist from a credentialed generalist, so expect to be pushed on the specifics of matters you actually owned: the practice area, the posture, the decisions that were yours. Verifiable specifics carry more weight than titles.
- Depth over breadth — one genuine specialization (corporate/transactional, litigation, regulatory and compliance, IP, employment and labor, or tax) probed several layers down.
- Evaluation judgment — whether you can articulate why an answer fails, not just that it does.
- Real seniority — partner, of counsel, counsel, senior associate, senior in-house, or GC, with ownership of matters. Clerkships and internships alone do not count.
- AI fluency — hands-on use of LLMs in professional work, and a calibrated sense of where they hallucinate.
Logistics
Full-time W-2 employment with Cincinnatus LLC (employer of record; payroll, benefits, and compliance run through them, not Mercor), with placement at a leading AI lab's extended workforce. 40 hours per week, initial engagement of six months. Work is remote day to day, but you must live in the Bay Area and be able to come on-site when required — on-site days are expected to be infrequent. Observed pay for comparable legal evaluation roles on the platform runs $50–90/hr; the band is not a guarantee and depends on practice area, seniority, and client scoping.