The work

You are handed the kind of material that fills an ordinary workday: a two-page memo, a client-facing deck, a budget model, a policy summary. Sometimes the model wrote it; sometimes you are comparing two model attempts side by side. Your job is to decide whether the artifact would actually survive being sent to a colleague or a client — is the arithmetic right, does the slide say what the title claims, is the conclusion supported by what precedes it, is anything quietly missing. Then you write down why, in prose a reviewer can audit.

Tasks arrive in batches with a rubric attached. Ratings are usually on a fixed scale with required written justification, and pairwise comparisons ask you to rank outputs and explain the deciding factor. Much of the difficulty is not in spotting a howler but in separating a genuine defect from a stylistic preference — polished, confident writing that is subtly wrong is the central failure mode you're being paid to catch.

What the screen looks for

Mercor runs an AI-led interview before any project match. It is checking that your degree and location claims are real, that your written English holds up under a follow-up question, and — most of all — that you can articulate the reason behind a judgment instead of reporting a verdict. Generalist listings attract volume, so the differentiator is almost always specificity: a concrete example of an error you caught, a rubric you followed, a case where you disagreed with a guideline and how you handled it. Expect probes across formats, including spreadsheet logic, since many candidates are strong on prose and weak on numbers.

Logistics

  • Fully remote, US or Canada residence required; work is asynchronous with no fixed shifts
  • Contract, hourly, paid per approved hour — observed rates for this listing sit in the $50–70/hr range, not guaranteed
  • Volume fluctuates by project; many contributors treat it as part-time alongside other work
  • No prior AI or ML background needed — guidelines and calibration materials are supplied
  • Expect an unpaid calibration or sample-task stage before live work begins