What the work involves
You build and judge evaluation material for frontier models on data science reasoning. In practice that means authoring problems a competent model should plausibly fail — a leakage-riddled feature pipeline, an A/B test with a peeking problem, a time-series split done wrong, a metric choice that flatters a useless classifier — then writing the reference solution and the rubric that separates a defensible answer from a fluent one.
Day-to-day tasks typically include:
- Writing prompts grounded in real analytical situations, with the data, code, or output a working analyst would actually see
- Grading multiple model responses against a rubric and explaining, in writing, exactly where reasoning breaks
- Reviewing other contributors' items for correctness, ambiguity, and unintended giveaways
- Occasional live comparison work: ranking outputs side by side and justifying the ordering
What micro1 screens for
The intake is an AI-led interview followed by task-based work samples. Screening looks for whether you have shipped analysis that someone depended on — not coursework, not tutorials. Expect follow-up questions that push past your first answer: which estimator, why that assumption, what you'd do if the residuals misbehaved. Reviewers also weigh writing quality, because a rubric that cannot be applied consistently by a second grader is worthless.
Statistical honesty matters more than tool breadth. Contributors who flag uncertainty, distinguish correlation from claim, and refuse to grade outside their competence tend to get more work.
Logistics
Fully remote and largely asynchronous, with tasks pulled from a queue against deadlines rather than fixed shifts. Most contributors commit 10–20 hours weekly, though volume fluctuates with project cycles and some engagements ramp sharply for short periods. Pay is hourly under an independent contractor agreement; the $100–200/hr band reflects observed rates and is not a guarantee for any given project or region.