What the work involves
You sit alongside a leading AI lab's research and program management teams, and your practice judgment is the deliverable. Day to day that means three overlapping streams: QA on legal knowledge-work tasks and model outputs, where you flag thin reasoning, missing steps, and answers that read fluently but would not survive a partner's or a regulator's review; authoring instruction specs and golden solutions that define what "correct" looks like on a given task; and designing harder evaluation sets and benchmarks that show whether the model is actually improving rather than just sounding better. A fourth stream is calibration — working with researchers and experts in adjacent domains to turn tacit legal judgment into explicit, teachable criteria that a non-lawyer annotator could apply consistently.
The tasks are built to mirror how the work is really done: not bar-exam trivia, but the reasoning behind a diligence memo, a discovery dispute, a compliance analysis, a claim construction, a withholding question. Written precision matters more than volume — feedback that says why an output fails and what the correct chain of reasoning was is the unit of value here.
What the screen looks for
- Verifiable seniority. JD from an accredited school, admission to at least one U.S. state bar in good standing, and 8–15 years of substantive post-qualification practice at a firm, in-house department, regulatory body, or court. Clerkships and internships alone do not count. If you are earlier in your career, or already at partner or GC level, apply to the matching seniority listing instead — the platform runs separate bands so rate tracks experience.
- Genuine specialization. Expect follow-ups that go a layer deeper than your résumé in one named practice area — corporate/transactional, litigation, regulatory and compliance, IP, employment and labor, or tax. Generalist answers tend not to survive the second question.
- Evaluation judgment. Can you tell a well-reasoned answer from a confident, plausible-sounding wrong one, and can you write down the distinction as a rule someone else could apply?
- Hands-on LLM use in your actual professional work, described specifically.
Logistics, stated plainly
This is not remote. It is a hybrid role based in the Bay Area, California, with required on-site days at the client's team multiple days a week. You must already live in the Bay Area or be willing to relocate at your own cost — no relocation assistance. The commitment is 40 hours per week for an initial six-month engagement, as a W-2 employee of Cincinnatus LLC, which serves as employer of record and administers payroll and benefits; Cincinnatus is a separate legal entity from Mercor. You'll work inside the client's own tools on client-issued accounts and equipment. The $85–120/hr band reflects rates observed on this listing and is not a guarantee — placement and rate depend on seniority, practice area, and the client's needs.