The AI training-data market is splitting in two. Generic annotation — labeling, transcription, simple preference ratings — is increasingly done by models themselves, and prices there keep falling. Expert evaluation — a physician grading clinical reasoning, a lawyer checking case-law application — keeps commanding higher rates.

Why the split

  • Model capability moved up the value chain. Frontier models handle everyday questions well; the remaining errors are the ones only experts catch.
  • Evaluation is the bottleneck, not generation. Labs can generate unlimited synthetic answers; deciding which are correct in a specialized domain still needs human judgment.
  • Reasoning models raised the bar. Grading a multi-step chain of reasoning in tax law or differential diagnosis is qualitatively harder than rating a chat reply.
  • Liability. In medical, legal, and financial domains, labs need documented expert review before shipping capabilities.

What it means for you

Platforms aren't hiring "annotators who know some finance" — they're hiring finance professionals who can explain why an answer is wrong, in writing, with structure. Expect screening to test exactly that, and expect the task mix to keep shifting toward rubric design and agent-trace grading, which favors senior professionals over generalists.