What the work involves
You build "worlds" — self-contained evaluation scenarios that put an AI model in the position of a public interest lawyer and then measure whether it behaves like one. In practice that means writing a fact pattern with real stakes (a Title VII pattern-or-practice claim, a NEPA adequacy challenge to an agency EIS, a CWA citizen suit against a permitted discharger, an asylum claim turning on changed country conditions), specifying the tools and records the model can reach, then drafting the reference output yourself: the brief section, the regulatory comment, the advocacy memo, the strategy note to a coalition partner.
The rubric is the part that carries the most weight. You are documenting how a senior public interest attorney distinguishes a defensible answer from a plausible one — standing and ripeness that actually hold up, remedies that a court can order, an administrative record argument that survives arbitrary-and-capricious review, an awareness of when litigation is the wrong instrument and a rulemaking comment or legislative campaign is the right one. Generic bar-exam recall is the failure mode you are explicitly writing against.
What the platform screens for
- Verifiable practice history. Mercor's screen probes named matters, your role in them, forum, posture, and outcome. Five-plus years at a legal aid organization, advocacy nonprofit, or government office (ACLU, Earthjustice, NRDC, Legal Aid Society, a state AG's office, or equivalent) is the baseline.
- Depth under follow-up. Expect to be pushed a level past your first answer — on exhaustion requirements, on the difference between facial and as-applied challenges, on what an EPA docket comment actually needs to preserve an issue for later review.
- Evaluation judgment. Can you say precisely why one model output is wrong rather than merely unpolished, and can you write that distinction as a rubric criterion someone else could apply consistently?
- Credentials. JD with active or prior bar admission, or an international equivalent, is strongly preferred. Prior rubric, training-material, or policy authorship helps.
Logistics
Fully remote and asynchronous. Contributors typically commit 10–20 hours per week with no fixed schedule, though onboarding usually includes a live calibration session and early submissions get reviewer feedback before you scale up volume. Pay is contract, hourly, invoiced through the platform; $90–100/hr reflects observed rates for this listing and is not a guarantee. Work is track-based — you can apply for the US track, the International track, or both.