The work

Each task is three artifacts. First, the design: you pick a hidden conclusion — a mechanism, a liability, a go/no-go call — and construct the forcing chain of evidence that makes that conclusion the only defensible one. Then the data room: assay tables, structures, patent excerpts, internal-style reports, literature figures, all traceable to real sources, and deliberately incomplete, indirect, or partly misleading. Finally the grader: a document that awards credit for the inferential steps a competent scientist would take, and withholds it for a right answer reached by luck or fluency.

Before submission you verify and calibrate your own task — confirming the chain actually closes, that no decisive value is unsourced, and that the problem isn't solvable by recall or a single lookup. Reviewers will challenge your difficulty calls and your science; you defend or revise. On the topics inside your task you are the final authority, which is also why the rework lands on you.

What the screen looks for

Mercor's screening is AI-led and follow-up heavy. It probes for concrete depth in at least one of mechanistic enzymology, structure-guided peptide or antibody design, fragment-based discovery, or structure-based medicinal chemistry — expect to be asked about specific assays you ran, specific structures you worked from, and where the data misled you. It also probes evaluation judgment: whether you can distinguish a memo that reasoned well from one that wrote well, and whether you can name what would make a problem trivially shortcuttable.

  • Naming real programs, targets, assay formats, and failure modes helps more than describing responsibilities
  • Expect a scenario where a model reaches your intended conclusion via wrong reasoning, and be ready to say how you'd score it

Logistics

  • Fully remote and asynchronous; you set your own hours
  • Minimum 20 hours per week, up to 40 — the floor is real, because tasks span design, build, calibration, and revision
  • Work happens in a browser-based studio plus Claude Code; a Claude Max subscription is required and fully reimbursed
  • Observed pay band $60–90/hr, not guaranteed; rates typically track demonstrated depth and task acceptance