What the work involves

You will be given documents — grant proposals and narratives, ESG disclosures, sustainability or materiality assessments, policy memos, contract summaries — either drafted by a model or drafted by you as a reference standard. Your job is to judge them the way a funder, regulator, or reviewing counsel would: is the eligibility logic correct, is the metric cited under the right framework, does the argument survive scrutiny, and would this document actually pass in the setting it imitates.

Feedback is structured, not prose commentary. Expect rubrics with dimensions like factual accuracy, completeness against a stated requirement, formatting convention, and reasoning quality. You score each dimension and then write a short justification that a second reviewer can follow without knowing your background. Some projects ask for the corrected version alongside the critique; others ask you to rank two model outputs and explain the margin between them. Tasks frequently involve Excel and PowerPoint artifacts, so comfort producing and critiquing those formats matters more than it sounds.

What the platform screens for

Mercor's screen is AI-led and follow-up heavy. It tests whether your claimed experience is specific — which grantmakers, which reporting standard, which stage of the process you personally owned — and whether you can articulate why a plausible-looking document is wrong rather than just marking it wrong. Vague seniority claims collapse quickly under probing. Reviewers who can name a concrete failure mode from their own work (a proposal rejected on an allowability rule, a disclosure flagged for the wrong scope boundary) do markedly better than those who describe responsibilities in the abstract.

Logistics

  • Fully remote, asynchronous, with per-task deadlines rather than fixed shifts
  • Typical commitments run 10–20 hours weekly; some projects offer more, and volume fluctuates between project cycles
  • You must be based in the US, Europe, UK, or Canada
  • Pay is hourly within an observed $80–160 range; the band reflects task complexity and domain scarcity and is not a guarantee