What the work actually involves
Each item arrives as a complete package: the task prompt as written, a transcript of what the AI agent did, the files it produced (documents, spreadsheets, code, schedules), and a scoresheet an AI grader already filled in. Your job is to audit that package, not to redo it. You answer four things — whether the task is one that genuinely comes up in your job described that way (realistic, borderline, or contrived); whether the domain content holds up (standards cited correctly, figures internally consistent, regulations current, constraints plausible); whether each of the grader's eleven yes/no judgments is right, item by item with the files open; and whether the final score is defensible against your own band of the work. Every judgment carries a short written reason, so the writing is a real part of the job, not an afterthought.
Budget about an hour per item. Most of that hour is spent inside the artifact: checking that a project schedule's float is actually consistent with its dependencies, that a set of financial statements ties, that a contract clause does what its recital claims, that code compiles conceptually and handles the case the task specified. The errors that matter are the subtle ones — a plausible-looking figure that doesn't reconcile, a regulation quoted in its superseded form, terminology used the way a layperson would use it rather than a practitioner.
What the screening looks for
The platform screen is field-specific and pushes on delivery history rather than titles. Expect to be asked what you personally produced and who relied on it — a close you owned, an opinion you signed, a plan other people worked to, code other engineers maintained after you left. Reviewers who only supervised the function, and candidates whose case rests on a certification without artifacts behind it, are explicitly a weaker fit. You will also be probed on grading judgment: can you disagree with a grader's call and say precisely why, without sliding into rewriting the work yourself.
Logistics
- Fully remote and asynchronous; items are claimed from a queue rather than scheduled.
- Roughly an hour per item, so the work fits alongside active practice.
- Observed band $70–85/hr, varying by field and demonstrated depth — stated as observed, not guaranteed.
- Requires near-native business English and enough spreadsheet/PDF fluency to check whether numbers reconcile.
- Apply to one of the five tracks — the one that matches what you actually deliver.