What the work involves
You receive AI-generated artifacts — a dedupe and merge plan for a Salesforce org, a lead-routing spec, a data dictionary, a field-mapping workbook for a HubSpot-to-Salesforce migration, an executive deck on pipeline hygiene — and grade them against a rubric. The judgment being captured is whether the output would survive contact with a real CRM: does the matching logic actually handle the false-positive cases, are the required fields and validation rules coherent, does the spreadsheet formula do what the narrative claims, is the governance step that a real RevOps team would insist on simply missing? You write structured written feedback naming the specific defect and why it matters operationally, not a general impression.
Task types rotate. Some batches are scoring passes with tight rubric dimensions; others ask for side-by-side preference judgments between two model outputs; others ask you to rewrite a section to demonstrate the correct answer. Slide-heavy batches are common, which is why proficiency in Google Slides and PowerPoint is a hard requirement rather than a formality — you are assessing layout, chart choice, and data-labeling accuracy alongside the substance.
What the platform screens for
- Verifiable operational history. Which CRM platforms, at what scale, in what role — record counts, integration surface, whether you owned the data model or consumed it.
- Depth under follow-up. Screens push past first answers into matching thresholds, survivorship rules, sync-conflict handling, and how you measured data quality rather than asserted it.
- Evaluation judgment. Whether you can separate a cosmetic flaw from a defect that breaks reporting, and whether you can defend a score against a plausible counterargument.
- Written clarity. Feedback that another reviewer could act on without a follow-up conversation.
Logistics
Fully remote and asynchronous, with work claimed from a queue rather than scheduled. Hourly and contract-based; the $80–120/hr band reflects rates observed on similar Mercor evaluation engagements and varies with credentials, task complexity, and batch — it is not a guarantee. Volume is genuinely uneven, so this suits people who want 5–20 hours a week of flexible overflow work rather than a stable full-time load. Expect a calibration exercise and rubric onboarding before paid batches begin.