What the work actually involves

You write the scenarios a frontier model gets tested on, then judge whether it handled them the way an experienced ops manager would. A task might be a mid-term endorsement with a retroactive effective date, a cancellation for non-payment where the notice timeline matters, a certificate request that asks for coverage the policy does not grant, or a renewal where the declarations page and the binder disagree. For each, you produce a "golden" reference response — processing instructions, a transaction review, an exception-handling plan, a service communication — at the quality you would accept from a senior processor, then grade model attempts against a structured rubric covering procedural accuracy, completeness, sequencing, and service standards.

The error patterns the lab cares most about are specific: wrong or impossible effective dates, missing signed applications or supporting documents, transactions posted out of order, reinstatements treated as if no lapse occurred, and — the big one — a model quietly making an underwriting decision when the correct answer was to refer it. Your written feedback explains not just that the output is wrong but which control it bypassed.

What the screen looks for

  • Chronology under pressure. Can you narrate a policy lifecycle in correct order and say what document or system state each step depends on?
  • The boundary discipline. Administrative execution versus underwriting, claims, legal, and actuarial judgment. Candidates who blur this do not pass.
  • System-level detail. Named policy administration platforms, workflow queues, transaction controls, QA sampling — concrete, not conceptual.
  • Writing that a research team can act on. Feedback has to be legible to a non-insurance reader while remaining technically exact.

Logistics

Fully remote and asynchronous, paid per completed task at an observed $800 rate that varies with task scope. Minimum 20 hours per week, with 40+ preferred; the role starts immediately and applications are reviewed on a rolling basis. Expect live onboarding office hours and periodic calibration sessions where graders reconcile scoring differences — these are usually scheduled during US business hours, so some overlap helps.