What the work involves

You are not writing documentation for a product — you are writing documentation problems that a strong AI agent should be able to solve, plus the rubric that decides whether it did. A typical task starts with you constructing a believable scenario: a firmware release with incomplete engineering notes, a medical device labeling change under regulatory pressure, an undocumented internal API with inconsistent code samples. You assemble the raw inputs (CSVs, PDFs, spreadsheets, spec fragments, SME interview notes), define the deliverable, then decompose "good" into 35+ discrete, checkable rubric items covering factual accuracy, precision of instruction, testability of steps, and fit for the stated audience.

The rubric is usually the hardest part and the part reviewers scrutinize most. Vague criteria ("clear and well-organized") get rejected; criteria that a second reviewer could apply to the same output and reach the same verdict get accepted. Expect asynchronous back-and-forth with project leads on early submissions as you calibrate to the benchmark's difficulty target and formatting spec.

What the platform screens for

  • A real portfolio in software, hardware, medical devices, or another regulated domain — 2–4+ years of shipped documentation, not adjacent content marketing.
  • Evidence you can write to a style guide and work in a documentation toolchain (structured authoring, docs-as-code, DITA, or similar).
  • Prior experience defining acceptance criteria, review checklists, or QA rubrics for technical deliverables.
  • Written English with genuine control of register — API reference voice versus end-user guide voice versus compliance language.
  • Judgment about failure modes: can you predict where a competent writer or model would get something subtly wrong?

Benchmark or evaluation experience is a nice-to-have, not a gate. Domain depth is the actual filter.

Logistics

Fully remote and asynchronous, contractor engagement. Compensation is output-based — paid per task that meets project spec — so effective hourly earnings depend on how quickly you produce accepted work; the $30–60/hr band reflects reported ranges rather than a guaranteed rate. A weekly minimum submission volume applies. micro1 typically fills these roles within 48 hours and expects first tasks within 24–48 hours of onboarding, so this suits people with immediate open capacity rather than a queued start.