What the work involves
You receive AI-generated work products that imitate the deliverables of a public-sector capture or proposal team: draft RFI narratives, RFP section responses, compliance and cross-reference matrices, past-performance write-ups, pricing and basis-of-estimate spreadsheets, and capability slide decks. Your job is to grade each artifact against a rubric and explain, in writing, exactly where it fails. Typical findings include invented FAR or DFARS citations, responses that ignore Section L instructions or misalign with Section M evaluation criteria, small-business subcontracting math that does not add up, missing certifications or representations, and slides whose formatting or claim substantiation would not survive a color team review.
- Score outputs on dimensions such as factual accuracy, responsiveness to the solicitation, compliance, and presentation quality
- Write structured critiques that a model trainer can act on — the specific defect, why it matters to an evaluator or contracting officer, and what a correct version would say
- Compare two candidate outputs and justify a preference when tasks are set up as pairwise comparisons
- Occasionally repair or rewrite a passage to demonstrate the standard you are describing
What the platform screens for
Mercor's screening is AI-led and interview-based. It probes whether your procurement experience is real and specific: which agencies or jurisdictions, which vehicles and contract types, whether you sat on the buy side or the response side, and what you personally owned on a given bid. Expect follow-up questions that go a layer deeper than your first answer. Reviewers also test evaluation judgment — whether you can separate a genuine compliance failure from a stylistic preference, and whether your written feedback is legible to someone without your background. Slides and spreadsheets matter here; the listing calls out proficiency in Google Slides and PowerPoint because a meaningful share of tasks are deck reviews.
Logistics
Fully remote and asynchronous. Work is drawn from a task queue with no fixed shifts, and most contributors treat it as part-time — commonly 10–20 hours a week, though volume fluctuates with project demand and can pause between cohorts. Hourly rates are set per project and per contributor; the $80–120/hr band reflects what has been observed for this listing rather than a guarantee. Reliable weekly availability and prompt responses to reviewer feedback tend to matter more than raw volume.