Evaluation pay is tiered by scarcity: how few people can do the task credibly. The moves that raise your rate all push you up that scarcity curve.

The tiers, roughly

Listed lowest to highest. The figures themselves live on the rates page on this site, which shows the current median and range by field, recalculated from live listings — the bands below are described in words because any number written here would be out of date within weeks.

  • Generalist annotation. Labeling, simple preference rating. Shrinking and increasingly automated.
  • Skilled evaluation. Coding review, STEM problems, multilingual reasoning. Needs demonstrable skill, not necessarily a license.
  • Credentialed review. Licensed clinical and legal work, senior finance. Verification required; demand is durable.
  • Design work. Rubric authoring, benchmark question writing, eval design. Usually offered to proven evaluators, above standard evaluation rates.

The moves that raise your rate

  • Verify everything. Upload licenses, certifications, and degrees — credentialed profiles are routed to better-paid queues.
  • Write superb critiques. Platforms grade your evaluations. Consistent, well-argued write-ups get you invited to rubric and design work — the top tier.
  • Specialize your profile. "Corporate lawyer, M&A, 8 years" beats "legal expert." Specific profiles match specific (better-paid) projects.
  • Run two platforms. Project waves are unsynchronized; a second pipeline smooths your hours more than any negotiation can.
  • Take the hard projects. Agent-trace audits and benchmark authoring pay above simple rating and teach you the skills the next tier screens for.

Realistic expectations

Task availability fluctuates with lab training cycles — treat weekly hours as variable and any single listing's rate as indicative. The evaluators who do best treat this as a professional side practice: reliable output, strict rubric adherence, responsive communication. Platforms notice, and routing follows.