What the work involves
You run adversarial sessions against conversational models and agents, then write up what broke and why. A typical shift mixes structured probing — working through an assigned taxonomy of misuse categories — with open-ended exploration where you follow a hunch across multiple turns to see whether a refusal can be eroded. Norwegian matters here because safety behavior often degrades outside English: a guardrail that holds in English may fail when the same request is framed in Norwegian, uses Norwegian cultural or legal context, or code-switches mid-conversation.
Outputs are artifacts, not opinions. You annotate failures, classify the vulnerability, rate severity, and produce attack cases another person can reproduce from your notes alone. Some batches cover sensitive material — bias, misinformation, harmful instruction-following. Topics are disclosed before you see content, higher-sensitivity queues are opt-in, and wellness resources are available.
What the platform screens for
Mercor's screen is AI-led and follow-up heavy. Expect it to test whether your red teaming experience is real: naming a technique is not enough, you will be asked to describe an attack you personally ran, what the model did, and why it worked. Native-level Norwegian is verified, and you should expect to be pushed on how safety failures manifest differently across the two languages. The screen also probes judgment — whether you can distinguish a genuine vulnerability from a model being unhelpfully cautious, and whether you write findings a customer engineer could act on.
Logistics
- Fully remote, asynchronous, project-based work assigned in batches
- Observed pay band of $48–62/hr, varying by project and specialty; not guaranteed
- Hours are flexible but batches carry deadlines; most contributors commit 10–20 hours weekly
- Assignments rotate across customers, so expect new taxonomies and guidelines periodically