What the work involves

You will be handed AI-generated artifacts from front-end revenue cycle workflows — parsed 271 responses, benefit summaries, coverage determinations, COB resolutions, prior-auth triggers — and asked to judge whether they are correct, complete, and defensible against the actual payer's rules. Much of the day is close reading: does the model's benefit summary match what the 271 loop actually returned, or did it hallucinate a copay tier? Did it catch that the patient has a Medicare Advantage plan rather than traditional Part B, and route the verification accordingly? You will write structured annotations explaining not just that an output is wrong but why a verification specialist would catch it, since the written rationale is what the training pipeline consumes.

Expect a mix of task types across a project: scoring individual outputs, comparing two model responses head-to-head, and occasionally authoring reference scenarios and SOPs that become gold-standard data. Volume and task design shift as the lab's priorities change.

What the platform screens for

  • Verifiable operational history — specific payers, clearinghouses, and EHRs you have worked in, and what your team's front-end denial rate looked like.
  • Depth under follow-up — the screen will push past the acronyms. Knowing 270/271 exists is table stakes; explaining what a service type code 30 response actually omits is the differentiator.
  • Evaluation judgment — whether you can separate "technically inaccurate" from "materially harmful," and articulate a consistent rubric rather than gut reactions.
  • Written clarity — annotations that a non-clinical ML engineer can act on.

CHAM or CHAA helps but does not substitute for hands-on payer work. Management experience matters mainly because it usually means you have owned the KPI and the root-cause analysis behind it.

Logistics

Fully remote and largely asynchronous, with work delivered through Mercor's task platform. Contributors commonly commit 10–20 hours per week, though some projects offer more; hours are self-scheduled against task deadlines rather than fixed shifts. Engagements are contract-based and scoped in waves, so continuity depends on the lab's roadmap. Occasional synchronous calibration calls with the research team are typical early in a project. Do not use real PHI in any submission — all work should use synthetic or fully de-identified scenarios.