What the work involves
You receive a set of prompts drawn from live Australian litigation and procedural practice — limitation periods, discovery and privilege, interlocutory relief, costs, jurisdiction and cross-vesting, pleading points, appellate avenues. You run each prompt through two LLM platforms, read both outputs closely, and score them against a standardized rubric rather than your own preferred grading scheme. Every evaluation ends with concise written feedback: what the model got wrong, why a practitioner would notice, and what the answer should have said. Volume matters less than consistency — the value of the dataset depends on you applying the same standard at hour nine as at hour one.
The hardest part is usually not spotting howlers but catching answers that are fluent and superficially correct while being wrong on the point that would actually decide the matter: citing a repealed rule, applying UCPR reasoning to a Federal Court proceeding, missing that a state's Civil Procedure Act differs materially, or inventing a case that reads exactly like a real one. Reviewers who verify citations rather than pattern-match to plausibility do best here.
What the platform screens for
- Provenance of your experience. The screen is built to distinguish firm and chambers litigators from adjacent roles. Expect direct questions about which courts, which side of the record, and what you personally ran versus supervised.
- Jurisdictional precision. Australian procedure is federated; screens probe whether you can name the applicable rules and differences between federal and state regimes without hedging.
- Rubric discipline. Interviews test whether you can score to someone else's criteria, including when the rubric produces a result you personally disagree with.
- Written concision. Feedback fields are short. Screens look for reviewers who can state a defect and its consequence in a few sentences.
Logistics
Fully remote and asynchronous — you choose your hours within the deadlines set for each batch. The initial commitment is roughly 10 hours, and reviewers who deliver consistent scoring are commonly invited into further batches. All work sits under NDA, including prompt content and platform identities. Pay in the $100–120/hr range has been observed for this brief; rates are set by the platform and vary with the engagement.