What the work involves
You are given a fact pattern or a contract excerpt and asked to do one of three things: write the gold-standard answer, grade two or more model responses against a rubric, or break a model that appears to be reasoning correctly. Typical tasks include assessing whether a model correctly distinguishes a condition precedent from a covenant, whether it applies the right measure of damages (expectation vs. reliance, consequential damages barred by an exclusion clause), whether it spots that a limitation-of-liability cap does not survive gross negligence in the governing jurisdiction, or whether its citation to a UCC provision or a state common-law rule is real and on point.
Much of the value you add is in written justification. A grade with no reasoning is unusable; annotators downstream need to know which step of the model's chain failed and why. Expect to write two to six paragraphs per item, in plain declarative prose, with pinpoint references to the contract language or authority you are relying on.
What the platform screens for
- Verifiable litigation experience — matters you actually staffed, not practice-area familiarity. Screeners follow up on procedural posture, who you represented, and what the disputed clause said.
- Doctrinal precision under follow-up. The AI interviewer will push on a doctrine you raise. Naming it is not enough; you need the elements and the jurisdictional variance.
- Calibration. Can you tell a wrong answer from an answer you would have written differently? Over-penalising stylistic divergence is the most common disqualifier.
- Hallucination detection. Whether you instinctively verify citations rather than trusting fluent prose.
Logistics
Fully remote and largely asynchronous — tasks are pulled from a queue with per-batch deadlines rather than fixed shifts. Most contributors commit 8–20 hours a week and can concentrate that on evenings or weekends, though calibration sessions and rubric-change briefings are occasionally scheduled live and skew toward US business hours. Work is engagement-based: volume fluctuates with project funding, and there is no guaranteed minimum. Pay bands are as observed on the platform and vary by task complexity and seniority tier.