AI companies spend billions training large language models — but compute and data alone don't teach a model to reason correctly about medicine, law, or finance. That takes structured feedback from people who actually practice those fields. That feedback loop is a labor market, and it's the one this site covers.
What the work actually involves
At its core, AI evaluation means reviewing AI-generated outputs and producing structured judgments. Depending on the project, you might:
- Rate an AI's answer to a finance question against a rubric
- Compare two AI coding solutions and explain which is better and why
- Flag a clinical explanation containing a dangerous error
- Audit a multi-step AI agent session and grade each decision
- Write hard questions in your specialty that expose shallow reasoning
You are not building AI systems and you need no machine-learning background. Your years of professional judgment are the product.
Why the market exists
Models learn from human feedback, but generic feedback produces generic models. Non-experts reliably reward confident-sounding wrong answers — which in specialized domains is worse than useless. So labs contract credentialed reviewers through platforms like Mercor and micro1, and pay a premium for them.
The practical shape of the work
- Fully remote and asynchronous — no meetings, no fixed hours
- Task-based — paid per hour or per task, tracked through the platform
- Flexible but variable — work arrives in project waves tied to lab training cycles
Most experts treat it as a side income of 5–20 hours per week alongside their practice.
What it pays
Rates are set per listing by the platform and its client, and they move as listings open and close, so this guide quotes no figures. The rates page on this site shows the current median and range for every field and platform, calculated from the live listings and recalculated on every visit. Two patterns hold across the board: generalist annotation sits at the bottom of the range, and licensed clinical and legal review, rubric design and benchmark authoring sit at the top. Treat any single listing's figure as indicative.
How to start
Browse the roles on this site, apply through the platform, pass its screening — on Mercor and micro1 that means a credential or skills check followed by an AI-led interview — and get matched to projects. Every role page here includes interview-prep questions prepared for that role — use them.