The work
You attack models on purpose. A typical session means picking a threat category from a provided taxonomy — misuse, disinformation, harassment, bias exploitation, agentic misbehavior — and building attack chains against a conversational model or agent until something breaks or you can credibly show it doesn't. Dutch matters here because safety behavior degrades across languages: a refusal that holds firmly in English often collapses in Dutch phrasing, idiom, or code-switching, and part of your job is finding exactly where that asymmetry lives.
When you find a failure, the deliverable is not the screenshot. It's a reproducible artifact: the prompt sequence, the conditions under which it fires, a classification against the project's vulnerability schema, a severity judgment, and enough annotation that an engineer who has never seen your conversation can replicate it. Reproducibility is the metric that separates useful red team data from anecdote.
Some projects touch sensitive content — bias, self-harm adjacent material, harmful instruction requests, misinformation. Mercor states that topics are communicated before exposure, higher-sensitivity work is opt-in, and wellness resources are available. Treat that as a real choice rather than a formality.
What the screen looks for
- Genuine native fluency in both languages, not conversational competence. Expect to be probed on Dutch register, regionalisms, and how you'd construct an attack that only works in Dutch.
- Prior adversarial experience with specifics: penetration testing, adversarial ML research, trust & safety investigation, disinformation analysis, or documented AI red teaming.
- Structured method over improvisation. Screens tend to reward candidates who reference frameworks, coverage thinking, and systematic variation over people who describe one clever jailbreak.
- Clear risk communication — can you explain severity to someone non-technical without inflating it.
Logistics
Fully remote and largely asynchronous. Work is contract and project-based, meaning volume fluctuates with customer demand; several contributors run this alongside other commitments at 10–25 hours per week, though full-time loads appear during active campaigns. Rates in the $48–62/hr band are what has been observed for this posting and vary by specialty depth and project — nothing here is guaranteed.