What the work actually involves
You build the material a model learns surety from. That means authoring realistic scenarios — a general contractor asking for a $12M single job on a $20M aggregate, a WIP schedule with suspicious underbillings, an obligee insisting on a non-standard performance bond form, a payment-bond claim from a second-tier supplier — and then writing the reference answer a seasoned underwriter or claims professional would give. After that you grade model attempts against structured rubrics covering credit judgment, capacity reasoning, bond-form accuracy, and default-resolution realism, and write feedback explaining precisely where the reasoning broke.
The failure modes you're hunting are specific: a model that treats a surety bond as first-party insurance, that pays a claim without touching indemnity or salvage, that reads a percentage-of-completion statement as if it were a manufacturer's balance sheet, that recommends takeover when tender or financing was obviously cheaper, or that confuses a Miller Act statutory bond with an AIA A312. Vague grades don't help the research team — the note has to name the error and the standard it violates.
What the platform screens for
Mercor's screen is AI-led and follow-up heavy. Expect it to take a claim like "I underwrote contract surety" and push: what aggregate did you carry authority for, what did your file look like when you declined, how did you set collateral. It is testing whether you can reason from a financial statement to a capacity number out loud, and whether you know where surety departs from property-casualty. Designations (AFSB, CPCU, CPA) help but do not substitute for being able to walk a WIP schedule.
Logistics
- Fully remote and largely asynchronous; you pick up tasks from a queue
- Minimum 20 hours per week, with 40+ preferred and generally better-supplied with work
- Live onboarding office hours and periodic calibration sessions — these are synchronous and attendance matters
- Rolling review, immediate start; pay observed at $800 per task, which varies by task length and complexity and is not guaranteed