What the work involves

You will spend most of your time in two modes. The first is authoring: constructing problems hard enough that a strong model gets them wrong or gets them right for the wrong reasons — a confounded A/B test readout, a leakage-prone feature pipeline, a survival analysis with informative censoring, a power calculation with clustered assignment — and writing the reference solution that makes the correct answer defensible. The second is grading: reading a model's response and deciding whether the statistical reasoning holds, then writing feedback precise enough that a trainer can act on it. "Wrong" is not a useful verdict; "applied a t-test to unit-level data when treatment was assigned at store level, so the standard errors are understated by roughly the design effect" is.

  • Design prompts and datasets in statistics, ML, causal inference, experimentation, and applied quantitative reasoning
  • Produce reference solutions with the assumptions and derivations shown, not just final numbers
  • Flag flawed methodology, silent assumption violations, misused tests, and confidently-worded nonsense
  • Write structured critiques against a rubric, and stay consistent across dozens of items

What the platform screens for

micro1's screening is AI-led and follow-up heavy. It will pick one technical claim from your background and push on it — which estimator, why that one, what you did when the assumption failed. Expect scenario items where two answers are both plausible and you must justify a ranking, and expect to be asked how you'd handle a model answer that reaches the right conclusion through invalid reasoning. Depth in one applied area beats a tour of every method you have heard of. The screen also verifies the concrete qualifications: recent industry experience, degree, location, and whether you can hold a technical explanation together in clear written English.

Logistics

Remote, contractor, and largely asynchronous — tasks are claimed from a queue rather than scheduled. Most contributors treat it as part-time alongside a primary role; sustained throughput matters more than any particular hours, and calibration sessions or rubric updates may occasionally be synchronous. Rates in the $245–280/hr band reflect what has been observed for this role and vary with track, seniority, and task type; they are not guaranteed.