What the work involves

You receive an artifact an AI model produced in response to a realistic engineering or manufacturing brief: a failure-analysis memo, an OEE improvement deck for a plant manager, a capex justification spreadsheet, a process capability summary, a supplier qualification report. Your job is to grade it against a rubric and write feedback that names the specific defect. That means catching the wrong unit conversion, the Cpk computed from an unstable process, the tolerance stack-up that ignores worst-case, the takt time that doesn't reconcile with the stated demand, the payback period that silently omits installation cost — and also the slide that buries the recommendation on page nine, or the chart with no axis units.

Most tasks mix three kinds of error. Factual and technical errors (wrong standard cited, misapplied formula, physically implausible number). Rigor errors (a conclusion the data doesn't support, an assumption never stated). Presentation errors (unreadable deck structure, broken spreadsheet formulas, inconsistent significant figures). Strong evaluators score all three separately rather than letting a polished-looking deck earn a pass on the engineering.

What the platform screens for

Mercor's screen is an AI-led interview that pushes on the specifics of your background: which processes, which line rates, which standards, what you personally owned versus what your team did. Expect follow-ups that get more concrete, not more abstract. It also tests whether you can criticise a document in writing — clearly, with a reason attached to each deduction — and whether you actually work fluently in Slides, Sheets, Excel and PowerPoint, since a large share of the artifacts are decks and workbooks.

  • 5+ years professional experience in engineering, manufacturing, or technical operations
  • Native or professional English fluency, with genuinely clean written prose
  • Hands-on proficiency with Google Slides / PowerPoint, Sheets / Excel
  • An advanced degree helps but is not required

Logistics

Fully remote and hourly. Work is asynchronous — tasks queue and you claim them, typically in blocks of a few hours rather than fixed shifts. Observed rates for this category sit around $80–120/hr, depending on domain depth and assessed feedback quality; nothing about pay or task volume is guaranteed, and queue depth fluctuates between projects. Many evaluators run this alongside consulting or a full-time role, committing 10–20 hours a week.