What the work involves
You'll review outputs an AI system produces when asked to do the analytics work you already do: build a denials scorecard, explain a swing in days in A/R, model net revenue realization under a payer mix shift, or recommend where to focus a cash acceleration effort. Tasks typically arrive as a prompt plus one or more candidate responses. You judge whether the metric definitions are right, whether the math holds, whether the data assumptions are stated, and whether the conclusion actually follows from the numbers — then write a structured rationale explaining your call.
The hard part is rarely spotting arithmetic errors. It's catching an AI that computes clean AR aging but silently mixes professional and facility populations, cites an HFMA MAP key with a subtly wrong denominator, treats a credit balance as collectible, or produces a confident root-cause narrative from a variance that's really a data quality artifact. Comparison and ranking tasks are common, as is rewriting a flawed output into a reference answer that shows the model what good looks like.
What the screen measures
- Verifiable specifics. Which reporting environment, what volume, which KPI set, what you personally owned versus what your team delivered.
- Depth under follow-up. Expect to define metrics precisely, name the denominator, and explain where the definition breaks down.
- Evaluation judgment. Whether you can separate "wrong" from "unsupported" from "technically correct but useless to an operator," and write that distinction down clearly.
- Written English. Rationales are the deliverable; terse or vague feedback fails calibration.
Logistics
Fully remote and asynchronous. Contributors commonly pick up 10–20 hours per week, though volume fluctuates with project batches and there is no guaranteed minimum. Expect an unpaid calibration exercise before paid work begins. The $100/hr figure reflects observed rates for this listing and is not a guarantee — Mercor sets rates per project and per contributor.