What the work actually is

Each item lands as a complete package: the task prompt as an AI agent received it, the agent's full working, the files it produced (usually a spreadsheet, memo, suitability write-up or client-facing letter), and a scoresheet an AI grader already filled in. You open the files and answer four things. Whether the task is one that would genuinely land on your desk, described the way a real requester would describe it — realistic, borderline, or contrived. Whether the domain content is right, meaning the standards, rate assumptions, regulatory citations, disclosure language and figures hold up under a practitioner's eye rather than a plausible-sounding reader's. Whether the grader's eleven binary calls — format, fabricated figures, requirements met, terminology, and so on — survive contact with the actual files, item by item, with your disagreements written out. And finally whether the score is defensible: you band the work yourself (could a practitioner use it? accept it as delivered?) and then compare the reported score to your band.

The last two questions are where most reviewers are weakest. Graders reward surface compliance — a model that has all the right tabs, a letter that contains all the required phrases. Your value is catching that the discount rate is applied to the wrong cash flow line, that a suitability rationale cites a rule that doesn't govern that product, or that a claims-servicing response promises a timeline the carrier can't meet. Equally, you have to be willing to say the grader was right and the work is fine, when it is.

What the platform screens for

The screen is AI-led and conversational, and it pushes on specifics. Expect to be asked which of the five tracks you're applying to and then interrogated inside it — a Series 7 holder gets asked about order handling and suitability files, an FP&A lead gets asked about close mechanics and forecast variance. Vague seniority claims collapse fast; named systems, named filings, named review responsibilities hold. The other thing it measures is whether you can criticise work without rewriting it: reviewers who instinctively start fixing the model instead of banding it are a poor fit, because the throughput assumption is roughly an hour per item and redoing the task blows that up.

US market grounding is a real filter, not a preference. The tasks turn on SEC, FINRA and state insurance specifics, so a strong non-US equivalent is generally not interchangeable here.

Logistics

  • Fully remote and asynchronous; items are claimed from a queue rather than scheduled.
  • Roughly an hour per item, with a short written rationale attached to every judgment — expect real typing, not clicking.
  • Observed band is $75–90/hr, varying by track and depth of credential; not guaranteed, and volume moves with project demand.
  • You need enough comfort with spreadsheets and PDFs to check that the numbers in front of you reconcile.