What the work actually is
You build cases, then judge answers to them. A task might start with a 58-year-old business owner considering a 1035 exchange out of a lapse-prone universal life policy, or a table-4 offer on a key-person case, or an FIA suitability review where the crediting caps and surrender schedule don't line up with the client's liquidity horizon. You write the fact pattern, produce a "golden" response at experienced-practitioner quality — the coverage recommendation, illustration comparison, underwriting assessment, replacement analysis or policy-service determination — and then grade model output against a rubric covering product accuracy, suitability, arithmetic and disclosure adequacy. Written feedback goes to a research team, so the reasoning behind a score matters as much as the score.
The failure modes you're hunting are specific: illustrations run at assumed rates presented as if guaranteed, MEC consequences ignored, indexed crediting described as market participation, confusion between a 1035 exchange and a taxable surrender, LIRP distributions modeled without regard to policy loan mechanics, or a replacement recommended with no comparison of surrender charges and new contestability. Models tend to be fluent and plausible here, which is exactly why the graders need to be people who have actually read an in-force ledger.
What the screen looks for
- Verifiable history. Two-plus years in life or annuity advisory, underwriting, new business, policy administration, product, or suitability review — with a named product family you can talk about in detail.
- Mechanics under follow-up. COI drag, cap and par-rate interaction, surrender charge schedules, rider costs and guarantee funding. Expect the AI screener to push a second and third layer on whatever you claim.
- Boundary judgment. Where product and suitability judgment stops and investment advice or formal tax and legal advice begins. Candidates who overreach here score badly.
- Writing. Feedback that a non-insurance researcher can act on.
Logistics
Remote and asynchronous, with onboarding office hours and periodic calibration sessions that are scheduled. Minimum 20 hours per week, with 40+ preferred; the role starts immediately and applications are reviewed on a rolling basis. Pay is observed at roughly $800 per task and varies with task complexity and length — treat it as a reported figure, not a guarantee. Designations (CLU, ChFC, FLMI, CFP) and an active or lapsed producer license or securities registration are bonuses, not gates.