What the work actually is
Each item arrives as a package: the task prompt an AI agent was given, a full transcript or record of what it did, the deliverables it produced (memos, spreadsheets, case plans, incident reviews, procurement documents), and a scoresheet an AI grader completed afterwards. You open the files and answer four things. Is this a task that would genuinely land on someone's desk in this job, described the way a supervisor would describe it? Is the domain content correct — the statutory citations, the reporting thresholds, the fee schedules, the mandated timelines — or is something subtly off in a way only a practitioner would spot? Do the grader's eleven yes/no findings hold up when you check them against the actual files? And when you band the work yourself (usable, acceptable with edits, not acceptable), does the reported score match your band?
You never perform the task. You do not rewrite the agent's memo or produce a correct version. The output is your judgment plus a short written reason for each call, which means the writing has to be tight and specific: "the grader marked 'no fabricated figures' as yes, but the $18,400 line item doesn't reconcile with the attached invoice total" rather than "figures look wrong."
What the screen looks for
- Operational, not analytical, experience. The platform is screening for people who have administered a program, carried a caseload, supervised sworn officers, run facilities or procurement, or issued findings someone was accountable to. Policy analysts who studied programs without running one are explicitly a weaker fit, as are private-sector equivalents — these tasks turn on public procurement rules, mandated reporting and public-records law.
- Specific citation habits. Expect follow-ups that push on where a rule comes from in your jurisdiction, what the actual deadline is, and what you do when a document is silent on something.
- Grader-disagreement judgment. Whether you can say a binary question was answered wrong and point to the file evidence, rather than substituting your general impression of quality.
- Numeric reconciliation. Enough comfort with a spreadsheet or PDF to check whether totals, headcounts, ratios and dates actually add up.
Logistics
Fully remote and asynchronous. Items are picked up from a queue and take about an hour each, so the work fits around a full-time public sector job — evenings and weekends are normal here, and most reviewers do a handful of items a week rather than a fixed shift. Observed pay is $50–55/hr; rates on Mercor vary by track, calibration performance and project, and are not guaranteed. Apply to the single occupational track that matches your actual title and years, not the closest-sounding one.