What the work involves

You work on the data layer behind coding models: authoring problems that a strong model should plausibly fail, producing reference solutions, and judging model output against real engineering standards rather than surface plausibility. A typical batch might ask you to build a self-contained repository bug with a hidden failure mode, write tests that distinguish a real fix from a patch that games the test, then compare two model attempts and explain in writing which is better and why. Other batches skew toward critique — reading a model's diff or its chain of reasoning and annotating exactly where it went wrong: silent type coercion, an unhandled concurrency case, a fix that passes tests but breaks an invariant.

  • Task authoring: original problems in your primary language, with tests, edge cases, and rubric notes
  • Response grading: pairwise preference or rubric scoring on model-generated code, with written justification
  • Failure annotation: pinpointing the first wrong step in a long reasoning or agentic trace
  • Review passes on other contributors' work, once you've built a quality track record

What the screen looks for

micro1's process is AI-led and fairly fast: an async interview where an interviewer probes your resume, then live technical work observed for reasoning rather than just a passing test. Expect follow-ups that go one layer deeper than your first answer — if you claim distributed systems experience, you'll be asked what actually broke in production and how you diagnosed it. The strongest signal is specificity: named systems, real constraints, decisions you'd defend. Writing matters as much as coding, because a grade without a legible rationale is unusable as training data.

Logistics

Fully remote, contractor, and mostly asynchronous — you claim work from a queue and deliver against per-task deadlines. Most contributors run 10–20 hours a week alongside a full-time job; some projects offer heavier sustained volume for a few weeks. Work is project-based, so expect gaps between engagements and rate variation across them. Pay bands are as observed and depend on the specific project, not guaranteed.