What the work involves
This is a talent network listing, not an active engagement. Mercor is assembling a bench of practicing data scientists so it can staff evaluation projects quickly when AI labs commission them. When work does arrive, it typically looks like judging whether an AI system can do real data science: reading a model-produced EDA notebook, statistical modeling write-up, feature engineering pipeline, or A/B test analysis, and scoring it against a rubric.
Day-to-day on a live project you would expect to:
- Write task-specific grading criteria before scoring — deciding what a correct causal identification strategy, a defensible model choice, or an honest uncertainty statement actually looks like for that task
- Score AI-generated (and sometimes human) deliverables against those criteria
- Write the justification, not just the number — reviewers care more about your reasoning trail than your score
- Absorb calibration feedback from senior reviewers and re-grade when your judgment drifts from the standard
What the platform screens for
Mercor's screen is AI-led and resume-anchored. It checks that your claimed experience is real and specific: which company, which team, what the dataset looked like, what you decided and why. Expect follow-ups that push past the summary — if you say you ran experimentation, you will be asked about power calculations, variance reduction, interference, or what you did when a test read positive for the wrong reason. The bar leans toward people who have shipped analyses at technology, research, or quantitative firms where the work was reviewed by other rigorous people.
The second thing measured is evaluation judgment. Grading is a different skill from doing: it rewards consistency, willingness to be explicit about criteria, and the discipline to score what is on the page rather than what you would have written. Vague, generous, or inconsistent scoring is the most common reason contributors wash out.
Logistics
Fully remote and asynchronous, contractor basis, project-scoped. Observed rates for this network are posted at $100–150/hr; actual rate depends on the specific project and your assessed level, and nothing is guaranteed. Because there is no immediate opening, applying is best understood as getting into the queue — some applicants hear back within weeks, others not for months. Live projects commonly run part-time at roughly 10–20 hours per week with flexible scheduling and hard turnaround windows on batches.