What the work involves
You are given prompts and model outputs drawn from real medicinal chemistry practice: a series of analogs with potency data and a question about which vector to pursue, a proposed retrosynthesis for a heteroaromatic core, a claim about metabolic soft spots or hERG liability, a docking-and-selectivity narrative for a kinase inhibitor. Your job is to decide whether the reasoning holds, mark where it breaks, and produce the answer a competent chemist would have written. Much of the value you add is in catching outputs that are chemically fluent but wrong — an implausible regiochemical outcome, an SAR conclusion that ignores confounded assay conditions, a bioisosteric swap that silently destroys a key hydrogen bond.
Assignments vary by project. Some batches are pairwise preference comparisons with written justification; others ask you to author original prompts with gold-standard solutions, or to red-team a model by constructing questions where surface-plausible answers fail. Expect calibration rounds against a rubric at the start, overlap scoring with other reviewers, and feedback cycles where your rationales themselves get reviewed.
What the screening looks for
- Verifiable depth. The AI interviewer probes specifics: what series you worked on, what the SAR actually showed, why a compound was killed. Vague answers unravel under follow-up.
- Judgment, not just knowledge. Can you say clearly why an answer is wrong, and how wrong, without hedging or over-flagging?
- Writing. Rationales must be legible to a non-chemist reviewer while remaining technically exact.
Logistics
Fully remote and largely asynchronous, with work claimed from a queue. Contributors commonly report 10–20 hours per week, though volume fluctuates with project cycles and can pause between batches. Engagement is contract-based, hourly, with rates in the $60–80 range observed for this specialty — not a guarantee, and often tiered by demonstrated calibration and task type.