What the work involves
You review model outputs that imitate journalism — wire-style briefs, explainers, interview summaries, headline sets, quote attributions — and judge them the way a desk editor would. Typical tasks include scoring two candidate drafts against a rubric and writing the rationale, flagging fabricated quotes or invented sources, checking whether a claim is actually supported by the linked material, and rewriting a passage to a gold-standard version that trainers can learn from. Some batches ask you to build prompts instead: construct realistic reporting scenarios (embargoes, anonymous sourcing, corrections, breaking-news uncertainty) that expose where models get things wrong.
The judgment being tested is editorial, not stylistic preference. Reviewers are expected to separate prose that reads well from prose that is verifiable, correctly hedged, and properly attributed — and to say plainly when a fluent draft is unpublishable. Written rationales matter as much as the scores; they are the actual training signal.
What the platform screens for
- Verifiable bylines or staff/stringer history at a named outlet, wire service, or broadcaster
- Depth under follow-up on a specific beat — politics, courts, business, health, science, sports, local government, foreign correspondence
- Consistency against a rubric on calibration samples, including cases where the pleasant-sounding answer is the wrong one
- Clear, fast written English; second-language reporting experience is a plus for localization batches
micro1's screening is AI-led: an asynchronous conversational interview probes your background and beat knowledge, followed by scored sample evaluations. Expect follow-ups that push past your first answer.
Logistics
Fully remote and asynchronous, with work drawn from queues rather than scheduled shifts. Most reviewers commit 10–20 hours a week; volume fluctuates with project cycles, so treat it as supplemental rather than steady. Contract engagement, hourly pay against approved time, with rates set per project and observed in the $40–75 range — higher end typically for niche beats, non-English work, or reviewers promoted to rubric-writing and quality audit.