What the work actually is
You are producing training and evaluation data for a frontier lab that wants models to reason about disability entitlement the way an experienced examiner does. A typical task starts with a scenario you build from scratch — a group LTD claim hitting the 24-month own-occupation transition, an IDI claim with residual earnings loss, a life waiver of premium file with a pre-existing-condition question — then a "golden" work product written at experienced-professional quality: the determination letter, the benefit calculation with offsets applied, the medical and vocational evidence summary, the transferable-skills analysis, or the appeal decision. After that you grade model attempts against a structured rubric and write feedback the research team can act on.
The failure modes you are hunting are specific: an any-occupation standard applied during the own-occupation period, functional conclusions asserted without an APS or FCE to support them, a missed SSDI or workers' compensation offset, elimination-period math that ignores a recurrent-disability provision, or a decision that reaches the right outcome while leaving the administrative record undocumented. Flagging that a model "sounds confident but is wrong" is not useful; naming the provision it misread and what the record would have had to contain is.
What the screen looks for
- Verifiable claims-side or underwriting history — carrier, TPA, employer program, or vocational practice — with at least one program type you can discuss in detail (group STD, group LTD, IDI, statutory DBL/TDI, or waiver of premium).
- Fluency with contract mechanics under follow-up: elimination periods, offsets and their coordination, pre-existing look-back and treatment-free windows, maximum benefit periods, residual and partial formulas.
- The discipline to stay in the adjudicator's lane — not diagnosing, not giving legal advice, not substituting your view for a treating provider's while still weighing the evidence they produced.
- Written analysis that a reviewer could defend on appeal, which is also what a rubric can score.
Logistics
Fully remote and asynchronous, with a 20-hour weekly minimum and a stated preference for 40+. Pay is observed at $800 per task rather than hourly, so task scope and turnaround matter more than clocked time; bands on Mercor vary by project and are not guaranteed. Expect onboarding office hours and periodic calibration sessions where graders reconcile scoring differences — attendance is part of the job, not optional extra.