What the work actually is
You open a finished item: the task prompt as an AI agent received it, everything the agent produced (documents, spreadsheets, method sheets, PO drafts, layout calcs), and the scoresheet an AI grader completed afterwards. Your job is audit, not production. You answer four things — whether the task is one that would genuinely land on a buyer's, supervisor's, or engineer's desk described that way; whether the standards, tolerances, regulations and figures inside the deliverable are correct and internally consistent; whether each of the grader's eleven yes/no calls survives contact with the open files; and whether the reported score matches the band you'd give the work yourself (usable by a practitioner, acceptable, or not). Every judgment carries a short written reason. Budget roughly an hour per item.
What the screen is looking for
Floor time, in a specific lane. There are five distinct roles — purchasing, first-line production supervision, mechanical engineering, industrial engineering, and shipping/receiving/inventory — and you should apply to the one that is genuinely yours rather than the one that sounds broadest. The signal is accountability: POs you actually placed, a shift whose schedule and people were yours, parts that got built, a time study that changed a line, a cycle-count variance you owned. Candidates whose experience is entirely inside planning software, or engineering that never left simulation, are described as a weaker fit. Expect follow-ups that push on numbers — a real reviewer can say what a PPAP element contains or why a receiving variance reconciles, not just that they've heard of it.
Logistics and pay
- Fully remote and asynchronous; items are picked up from a queue rather than scheduled.
- About one hour per item, so throughput is naturally low-volume and fits alongside a full-time manufacturing job.
- Observed band is $55–65/hr, not a guarantee — rates on Mercor vary by lane, seniority and calibration performance.
- You need near-native business English for the written rationales, and enough spreadsheet/PDF fluency to check whether figures in a file reconcile.
Calibration matters more than speed here. Reviewers who flag a grader error and explain the practitioner consequence in two sentences are worth more to the program than reviewers who agree with every checkbox.