What the work actually involves

You write the material a frontier model gets trained and tested against. That means constructing scenarios that a working reinsurance professional would recognize as real — a quota share with a sliding-scale commission and a loss corridor, a facultative submission with awkward attachment relative to the underlying program, an MGA binding authority where the fronting carrier's retention and collateral arrangement determine who actually carries the risk. For each, you produce a "golden" reference response at the quality you'd expect from an experienced practitioner, then grade model attempts against a structured rubric covering structural accuracy, contract logic, financial consequence, stakeholder roles, and market realism.

The grading is where most of the judgment sits. Models are fluent about reinsurance and frequently wrong in specific, diagnosable ways: collapsing attachment and exhaustion, treating ceding commission as a cost with no earnings mechanics, confusing a fronting carrier's contractual liability with its net economic position, assuming claims control follows the paper rather than the wording, ignoring collateral and counterparty exposure entirely, or answering a treaty-structuring question as though it were a direct underwriting or claims-adjudication question. Your written feedback has to name the error and explain the consequence, not just mark it down.

What the screening looks for

The intake is AI-led and probes depth by follow-up rather than by credential list. Expect to be pushed on one structure you claim to know until the answers get specific — actual layer arithmetic, actual bordereaux and recoverables mechanics, actual descriptions of who owes what to whom in a fronted program. Generalist insurance framing tends to surface quickly. Designations (CPCU, ARe, ARM, ACAS, FCAS) help but do not substitute for being able to walk a structure end to end. Both-sides experience — ceded and assumed, or carrier and MGA — is treated as a strong signal because it predicts catching one-sided reasoning in model output.

Logistics

  • Fully remote and asynchronous; work is drawn from a task queue rather than assigned shifts.
  • Minimum 20 hours per week, with 40+ preferred; the listing states the role starts immediately with rolling review.
  • Live components are limited to onboarding office hours and periodic calibration sessions, which matter for rubric consistency.
  • Pay is observed at $800 per task and is not a guarantee; task scope and volume vary with the lab's needs.