What the work involves

You build the material AI models are tested against. A typical task starts with a realistic enterprise insurance scenario you construct — a complex commercial property submission with layered coverage and questionable loss history, a multi-party liability claim with reservation-of-rights considerations, a reserving review that changes a quarterly picture, or a state filing objection cycle. You then write the reference output a senior operator would actually produce (an underwriting rationale, a claims strategy memo, a reserving narrative for leadership) and a rubric that spells out what separates real judgment from a fluent but hollow answer.

The rubric is usually the hardest and most valuable part. Models can recite NAIC terminology, ISO form names, and textbook loss-reserving definitions. What they miss is the practitioner reasoning: when a facultative placement makes sense versus treaty capacity, why a claims file escalates before a legal demand arrives, how a pricing indication changes under credibility weighting, what an examiner actually flags. Your rubrics need to make those distinctions gradable by someone who is not you.

What the platform screens for

  • Verifiable, specific experience — lines of business, portfolio size or premium volume, systems you personally worked in (Guidewire, Duck Creek, policy admin stacks, SAS/R actuarial environments), and what you owned rather than observed.
  • Depth under follow-up — screens ask a general question, then push on the second and third layer. Vague framework recitation reads as thin quickly.
  • Evaluation judgment — can you look at two plausible model answers and articulate precisely why one reflects real carrier practice and the other does not.
  • Writing — reference outputs and rubrics are documents other reviewers rely on; clarity matters as much as correctness.

Logistics

Fully remote and asynchronous. Contributors typically commit 10–20 hours a week, self-scheduled, with work assigned in batches; some contributors take more when volume is available and less when it isn't. Engagements are contract-based and continuity depends on the project pipeline and calibration scores rather than a fixed term. Pay in the $50–60/hr range has been observed for this listing; actual rates vary with specialization and are set by the platform.