What the work actually involves

You'll spend most of your time reading model responses to hard prompts and deciding whether they should have been produced at all — and if so, whether the framing, hedging, and factual content were right. Tasks arrive in batches: sometimes side-by-side comparisons where you rank two completions and justify the ranking, sometimes single-response policy adjudication against a rubric written by a lab's safety team, sometimes red-team style prompting where you probe for failure modes in a specific hazard category. Expect content spanning misinformation and political persuasion, self-harm and crisis language, violence, cyber-offensive requests, and dual-use bio and chem questions. The hardest calls are rarely the obvious refusals; they're the requests where a legitimate user need sits next to a plausible misuse path, and your written rationale has to explain where you drew the line and why.

Rationale quality matters more than throughput here. Labs use these annotations to train and to calibrate their own policies, so a defensible three-sentence explanation of a borderline call is worth more than a fast label. You may also be asked to flag rubric gaps — cases the policy doesn't cleanly cover — which is one of the more valued contributions on safety projects.

What the screen looks for

  • Verifiable professional history. Five-plus years in AI safety, trust & safety, content policy, journalism, public policy, law, security research, or a lab science. Mercor's interview will push on specifics: what you adjudicated, under whose policy, at what volume.
  • Domain depth under follow-up. Generalist safety intuition doesn't survive the second question. Named hazard-area familiarity — crisis intervention protocols, election integrity policy, CBRN dual-use norms, vulnerability disclosure — is what separates candidates.
  • Evaluation judgment. Can you apply a rubric you disagree with, consistently, and log the disagreement separately? Can you distinguish an unsafe response from a merely unhelpful one?
  • Writing. Clear, compact English under time pressure. Rationales are the deliverable.

Logistics

Fully remote and largely asynchronous, with work claimed from a task queue rather than scheduled shifts. Contributors commonly report 10–25 hours a week, though availability fluctuates with client demand and specific hazard-area needs — some weeks the queue for your category is thin. Engagements are contract-based and hourly, with the $60–70/hr band reflecting observed rates on this listing rather than a guarantee; actual offers vary by project, seniority, and specialization. The material is genuinely difficult to read at volume, and anyone considering this should weigh sustained exposure to self-harm and violent content realistically.