The work
You build the material a frontier model learns claims reasoning from. In practice that means constructing fact patterns with real friction — a reservation of rights that may have been issued too late, excess layers arguing over attachment, multiple insureds with conflicting interests, allocation across policy periods, a mediation where the demand exceeds granted authority — then writing the "golden" response a strong technical claims director would produce. Work products mirror what you already write: coverage-position analyses, claim strategies, authority requests, litigation-management plans, reserve rationales, escalation memos.
The other half is grading. You score model outputs against structured rubrics and, more importantly, explain in writing why a position fails: it cites a policy provision that doesn't say what the model claims, it recommends a denial the record can't support, it settles above authority without escalating, it drifts into rendering legal opinions a non-lawyer claims professional shouldn't render. That last category matters a great deal to this project — recognising the boundary between an operational coverage position and formal legal advice is a graded skill here, not a footnote.
What the screen looks for
- Concrete claim history: lines handled, severity range, whether you owned coverage positions or only liability, whether you held or requested authority.
- Whether you can reason from policy language rather than from habit — insuring agreement, exclusions, conditions, endorsements, and how they interact across layers.
- Discipline with incomplete records: separating established facts, assumptions, open investigative items, and recommended next steps without collapsing them into one confident paragraph.
- Escalation judgment: when counsel, a coverage opinion, or home-office review is genuinely required versus when a claims professional decides.
- Writing that holds up under file review. Rubric grading is a writing job.
Logistics
Fully remote and largely asynchronous, with onboarding office hours and periodic calibration sessions you're expected to attend. Minimum 20 hours per week, with 40+ preferred; the listing states $800 per task, and observed pay on Mercor projects varies with task complexity and calibration performance rather than being fixed. Starts immediately, rolling review.