What the work involves
You receive AI-generated artifacts that imitate real pharma and clinical deliverables: a study synopsis, a mechanism-of-action deck, a safety summary table, an investigator brochure section, a payer value dossier slide, a competitive landscape spreadsheet. Your job is to score each one against a rubric and write feedback that a reviewer who is not a subject-matter expert could act on. That means separating three failure classes — factual and scientific errors (wrong endpoint, misattributed trial, implausible pharmacokinetics, misread inclusion criteria), rigor errors (unsupported causal claims, missing comparator, statistics that don't follow from the data), and presentation errors (broken slide hierarchy, unreadable figures, inconsistent units, malformed tables).
Expect to cite your reasoning. "The AUC figure is wrong" is not a usable evaluation; "the stated AUC is inconsistent with the dosing regimen on slide 4, and the label reports X" is. Task volume and artifact type shift as projects rotate, so comfort with ambiguity and a willingness to read a rubric closely both matter more than speed.
What the platform screens for
Mercor's screening is AI-led and follow-up heavy. It probes whether your stated experience holds up under specifics — which therapeutic areas, which phase, which regulatory pathway, what you personally produced versus reviewed. Generic answers get drilled into. The screen also tests evaluation judgment directly: whether you can distinguish a confidently wrong output from a merely incomplete one, whether you flag hallucinated citations, and whether you can hold a consistent standard across similar submissions. Slides proficiency is checked in practice, not just claimed — deck critique is a real part of the work.
Logistics
- Fully remote, contract, hourly. Observed rates for this role band $80–120/hr; actual offers vary by project and expertise and are never guaranteed.
- Asynchronous — no set shifts. Most contributors work in blocks of a few hours, with expectations set per project.
- Part-time volume is typical; sustained availability tends to matter more than a large weekly commitment.
- Native or professional English fluency is required because the deliverable is written feedback.