What the work involves

The core task is epidemiological judgment applied to AI output. You will be handed patient-population estimates — sometimes model-generated, sometimes drawn from analyst-style write-ups — for a specific drug or indication, and asked to determine whether the estimate holds up. That means walking the funnel: prevalence or incidence source and vintage, diagnosis rates, treatment rates, line-of-therapy restrictions, biomarker or subpopulation gating, and what actually remains as an addressable population. You will flag double-counting across indications, silent extrapolation from one geography to another, prevalence estimates applied to incident-only settings, and claims-based denominators that quietly exclude uninsured or Medicaid populations.

The second half of the work is rubric authoring and critique. You write the criteria another evaluator — or a model — should use to score a population estimate: what counts as adequate sourcing, when a range is more honest than a point estimate, how much weight a missing sensitivity analysis should cost. You will also review rubrics written by others and argue when they reward surface-level citation over defensible method.

What the platform screens for

  • Real fluency with claims, EHR, and registry data — specific datasets you have used, and the known biases of each
  • Ability to reconstruct a diagnosed → treated → addressable funnel out loud and name the assumption at each step
  • Judgment under ambiguity: whether you can say an estimate is wrong and explain by roughly how much and in which direction
  • Familiarity with how pharma and biotech teams actually use these numbers in program decisions
  • Written clarity, since most deliverables are prose critiques and rubric text rather than code or tables

Logistics

Fully remote, US-based only, and asynchronous — you pick up batches on your own schedule with per-batch deadlines rather than fixed hours. Expect 10–20 hours total across a one- to two-week pilot. Rates in the $170–270/hr range have been observed for this listing; actual offers depend on screening outcome and scope. Strong contributors are typically invited into recurring work as the evaluation set expands.