What the work involves

You receive an AI-generated artifact and the prompt or brief that produced it: a quarterly performance readout, a cohort analysis in Sheets, a regression summary, a deck translating a dataset into recommendations. Your job is to determine whether the analysis actually holds up. That means checking whether the right metric was chosen, whether the denominators and date windows are consistent, whether a correlation has been quietly upgraded to causation, whether the chart type suits the data, and whether the headline on the slide is supported by the numbers below it. You then write structured feedback tied to the rubric — specific enough that a model trainer or a reviewer can act on it without re-deriving your reasoning.

Most tasks combine a numeric score across several rubric dimensions with a written rationale. Expect to open and audit the underlying spreadsheet, not just read the summary. Common findings include silent unit errors, misapplied statistical tests, unlabeled axes, survivorship in the sample, and confident narrative built on a small or biased slice of data.

What the platform screens for

  • Verifiable depth: five or more years doing quantitative analysis professionally, with concrete examples of readouts you owned and who consumed them.
  • Statistical judgment under follow-up: the AI interviewer will probe an answer two or three levels deep, so precision beats breadth.
  • Tooling fluency: real comfort in Excel and Google Sheets, plus the ability to critique slide construction in PowerPoint or Google Slides.
  • Writing quality: feedback that is specific, calibrated, and free of hedging.
  • An advanced degree is a plus but does not substitute for applied experience.

Logistics

Fully remote and asynchronous. Work is drawn from a task queue; most contributors log 10–25 hours a week, though volume fluctuates with project demand and can pause between engagements. Rates in the $80–120/hr band are what has been observed on comparable Mercor evaluation work and are set per project, not guaranteed. Onboarding typically involves a rubric calibration exercise before paid tasks begin.