The work

This is comparative benchmarking, not content writing. You receive a set of prompts covering realistic UK legal questions — most weighted toward commercial and contract matters — and run each one through two separate LLM platforms. You then read both answers closely, decide where each one is correct, incomplete, or confidently wrong, and score them against a rubric the client supplies. Every evaluation ends with a short written justification: what the model got right, what it missed, and what a client would have been misled about if they had relied on the answer.

The hard part is calibration. A response that cites the right statute but misstates its application is not the same failure as one that invents authority, and the rubric expects you to distinguish them consistently across dozens of items. Reviewers who do well here treat the rubric as binding rather than advisory, and write feedback that another lawyer could audit.

What the screen looks for

  • Verifiable UK in-house experience — five or more years advising a business from inside it, with specifics on sectors, deal types, and the kinds of questions that actually crossed your desk.
  • Depth under follow-up — the AI interviewer will push past your first answer on commercial and contract points, so surface-level familiarity tends to show quickly.
  • Evaluation judgment — how you reason about a plausible-sounding but flawed answer, and whether you can separate style from substance.
  • Precision in written English — your feedback is the deliverable; ambiguous prose is a real disqualifier.

Logistics

Fully remote and asynchronous. The initial commitment is around 10 hours, and Mercor has indicated strong potential for additional volume if the benchmark continues. All work is covered by NDA. Pay is stated by the platform as USD 150–160 per hour — observed for this listing, not a guarantee, and final rates depend on the platform's own assessment. Work is done on your own schedule against submission deadlines rather than fixed hours.