What the work involves
You receive artifacts an AI model produced in response to a healthcare operations prompt: a spreadsheet modeling OR block utilization, a memo on ED boarding reduction, a slide deck presenting patient throughput metrics to a hospital executive committee. Your job is to grade it the way you would grade a junior analyst's work before it went in front of a CNO or COO — checking whether the numbers reconcile, whether the operational logic holds, whether the recommendations are actionable in a real hospital with real constraints, and whether the formatting would survive a board meeting.
Most tasks pair a rubric with a written response. You score dimensions like factual accuracy, analytical rigor, completeness, and presentation quality, then write structured commentary explaining each deduction with specifics. Vague feedback ("the analysis is weak") gets flagged in QA; useful feedback names the error, explains why it matters operationally, and points to what a correct version would look like. Expect to catch things like invented benchmark figures, length-of-stay calculations that ignore observation status, staffing ratios that violate state minimums, or a deck whose chart axes contradict its headline claim.
What the platform screens for
- Verifiable operational experience. Five or more years actually running or analyzing healthcare operations — throughput, capacity, staffing, revenue cycle, perioperative services, ambulatory access, supply chain. Screens probe for specifics: which systems, which metrics you owned, what changed.
- Depth under follow-up. Mercor's AI interview asks a general question, then narrows. Saying you improved OR utilization invites questions about your baseline definition, first-case on-time start rates, and how you handled block release policy.
- Evaluation judgment. Can you separate a stylistic preference from a substantive error, and calibrate consistently across many samples rather than drifting harsher or softer?
- Deck and spreadsheet fluency. Slides and Excel/Sheets proficiency is a hard requirement because a large share of artifacts are presentations and models.
Logistics
Fully remote and asynchronous. Work is claimed from a queue with deadlines rather than scheduled shifts, so hours are flexible but genuinely variable — volume comes in waves tied to project cycles. Most contributors treat this as part-time alongside a primary role; 10–20 hours a week is common when work is available. Pay is hourly, invoiced through Mercor. The $80–120/hr band reflects rates observed on comparable healthcare evaluation projects and is not a guarantee for any given assignment.