What the work involves

You write questions a language model should be able to answer if it genuinely understands biology — and then you grade what it produces. Tasks vary by project: authoring graduate-level prompts with defensible reference answers, comparing two model responses and writing a rationale for which is stronger, flagging fabricated citations or invented mechanisms, or annotating where a model's reasoning chain breaks down even when its final answer happens to be right. Some batches lean toward experimental interpretation (here is a figure, here is a claim — is the claim supported?), others toward factual precision in areas like pathway regulation, immunology, or population genetics.

The grading rubric is usually stricter than a journal reviewer's, because the goal is to separate plausible-sounding text from correct science. Reviewers who write "looks fine" get filtered out quickly; the value you add is in the specificity of your critique — naming the wrong enzyme, the misapplied control, the conflated timescale.

What the platform screens for

  • Verifiable graduate training in a life-science discipline (PhD, or MSc plus research output) and a subfield you can be pressed on
  • Ability to state a claim, cite the basis for it, and hold up under two or three rounds of follow-up questioning
  • Written reasoning that a second expert could audit — not just verdicts
  • Comfort saying "the literature is unsettled here" rather than inventing certainty

micro1's screening is AI-led: a structured interview with adaptive follow-ups, generally followed by a paid or unpaid sample task in your stated area.

Logistics

Fully remote and asynchronous. Contributors typically commit 10–20 hours per week, though volume fluctuates with project cycles and can pause between batches. Work is contractor-based, invoiced hourly or per completed task depending on the project. Expect an initial calibration period where your first submissions are reviewed closely before you're given higher-trust or higher-rate queues.