What the work involves
This is production-grade engineering applied to model training rather than shipping features. On a typical project you might author a non-trivial task in your primary language (a bug to localize, a refactor to perform, an API to implement), write the reference solution and test suite, then run a model against it and document precisely where and why the output fails. Other weeks skew toward evaluation: comparing two model attempts at the same pull request, ranking them against a rubric, and writing the justification that teaches the reward model what "correct" means beyond passing tests — readability, edge-case handling, idiomatic style, security posture.
Expect a mix of:
- Task authoring: repo-grounded problems with deterministic tests and clear acceptance criteria
- Output critique: line-level annotations on model diffs, explaining root causes rather than restating symptoms
- Preference comparison: side-by-side ranking with written rationale
- Rubric feedback: flagging where the grading spec is ambiguous or gameable
What the platform screens for
micro1's intake is AI-led — a structured async interview covering your background, then live technical depth with follow-up probing. Screeners are looking for verifiable specifics: which systems you owned, what you'd do differently, how you reason about a failure you can't reproduce. Vague seniority claims don't survive follow-up. Beyond coding ability, the bar is writing: your critiques are the training signal, so terse-but-complete English prose matters as much as your ability to spot the bug. Take-home or live coding in your stated primary language is standard, and self-reported expertise is verified against it.
Logistics
Fully remote and largely asynchronous, with occasional overlap calls for calibration or project kickoff. Commitments range from roughly 10 hours a week to full-time; most contributors set their own hours against weekly task quotas or throughput targets. Work is contract-based and project-scoped, so volume fluctuates with client demand — treat it as variable rather than guaranteed, and expect a paid calibration round before full-rate work begins.