What the work involves

You'll be assigned to one of two task streams. Authoring means building original questions in your subdomain — pharmaceutical manufacturing, industrial and synthetic biology, medical research and drug discovery, or agricultural, environmental and food biology — that test reasoning rather than recall. Each item needs a self-contained stem with no missing assumptions, one defensible correct answer, nine plausible distractors that a competent PhD could fall for, a markdown Chain-of-Thought solution with clean intermediate steps, and one to five references from peer-reviewed journals or university repositories. You also assign a difficulty tier: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above).

Verification is the mirror image: you receive pre-written questions and interrogate them for ambiguity, incompleteness, imprecision, or outright unsolvability. You edit where the item can be saved, flag it where it can't, re-rate difficulty, and document what you changed and why. Verification tends to be less glamorous and more forensic — a great deal of the value is in catching the item where two of the ten options are technically both correct.

What the platform screens for

Mercor's screening is AI-led and leans on verifiable specificity. Expect probing on your actual subdomain — not "biology" but the assays, regulatory constraints, pathways, or organisms you personally worked with — with follow-ups that go a layer deeper than your first answer. Screens also test item-writing judgment: whether you can distinguish a hard question from an ambiguous one, whether you understand what makes a distractor plausible rather than merely wrong, and whether you can justify a difficulty rating against a defined rubric. Credentials matter (PhD or doctoral candidate; master's accepted with exceptional depth), but demonstrated depth under follow-up matters more.

Logistics

  • Fully remote, fully asynchronous — no fixed hours or standing calls
  • Expected commitment 10+ hours/week; throughput expectations are per-item, not per-hour
  • Observed pay band $60–75/hr, varying by domain and task type; not guaranteed
  • Work is typically batched, so volume can fluctuate between project waves