We track public listings across the expert data platforms. What stands out this quarter:

Role types on the rise

  • Agent-trajectory evaluation. Reviewing multi-step agent sessions — tool calls, browsing, code execution — and grading whether each step was justified. Barely existed as a category a year ago.
  • Rubric authoring. Platforms increasingly want experts to write the grading criteria, not just apply them. Pays above simple evaluation.
  • Domain red-teaming. Probing models for confident-but-wrong answers in medicine, law, and finance.
  • Non-English expert review. Not translation — evaluation of model reasoning in the reviewer's native language. Japanese, Korean, Arabic, German, and Hindi lead demand.

Screening trends

  • Credential verification is stricter than in 2024–25 — licenses and bar membership checked more often.
  • Written critique samples are near-universal: read an AI answer in your field, identify the errors, explain them, timed.
  • AI-led interviews are the default first pass on both major platforms.

Pay observations

Rates are published per listing and move as listings open and close, so this piece quotes none. The rates page on this site shows the current medians and ranges by field and by platform, recalculated from live listings. Licensed clinical and legal review and rubric-design work sit at the upper end of those ranges. Availability still fluctuates project by project — treat single listings' rates as indicative.