What the work involves
You spend your sessions attacking models rather than assisting them. A typical task queue mixes structured probing — working through a taxonomy of harm categories, running scripted attack patterns, and checking whether a model holds its guardrails across a long multi-turn conversation — with open-ended adversarial exploration where you invent a novel jailbreak, role-play framing, or prompt injection and see how far it goes. Language matters here: many safety filters are trained and tuned primarily on English, and attacks that route through European, African, or Asian varieties of Portuguese, code-switching, or translation framing often succeed where the English equivalent fails. Finding and documenting those gaps is a large part of the value you add.
Every finding needs to survive someone else reproducing it. You write up the attack case with the exact prompts, the model's response, the harm classification under the project taxonomy, and a plain-language explanation of why it matters. Reports that read as "the model said something bad" without a reproducible path and a severity rationale tend to get rejected. Work touches bias, misinformation, harassment, and other harmful-behavior categories; topics are communicated before exposure, higher-sensitivity queues are opt-in, and wellness resources are provided.
What the screening looks for
Mercor's process is AI-led: a résumé and profile parse, then a recorded conversational interview, then paid or unpaid sample tasks. The interview probes whether your red teaming experience is real and specific — expect follow-ups asking you to name a framework you've used, walk through an attack you built, or explain why a particular jailbreak worked mechanically. Generic security vocabulary without a concrete case behind it reads poorly. Bilingual fluency is verified conversationally and through written samples; the exclusion of Brazilian Portuguese is a sourcing constraint on this project, not a judgment of the variety, so be explicit about which variety you speak natively.
Logistics
- Fully remote and asynchronous; you claim tasks from a queue rather than working fixed shifts.
- Contributors commonly report 10–30 hours per week, with volume fluctuating by customer project.
- Observed rate is $29–45/hr, positioned by experience depth and language coverage — not guaranteed, and subject to project.
- Expect assignment across multiple customers over time; adaptability to new taxonomies and playbooks is part of the job.