What the work involves

You attack models on purpose. A typical session means picking a threat category from a provided taxonomy — jailbreaks, indirect prompt injection, misuse elicitation, bias exploitation, multi-turn social engineering — and running structured adversarial conversations against a target model or agent in English, Danish, or across the two. Danish matters here because safety guardrails are usually tuned on English data; refusals, slang, euphemism, and cultural framing often behave differently once you switch languages or code-switch mid-conversation. When something breaks, you annotate the failure, classify the vulnerability, and write it up so someone else can reproduce it from your notes alone.

The output is data, not vibes. Attack cases, labeled transcripts, severity judgments, and short reports go to customers who use them to patch systems. Consistency is graded as heavily as creativity: two experts working the same taxonomy should classify the same failure the same way.

What the screen looks for

  • Verifiable red teaming history — adversarial ML work, penetration testing, abuse/trust-and-safety analysis, or socio-technical probing. Expect follow-ups asking what you actually did, not what your team did.
  • Native-level Danish and English, with the ability to construct adversarial prompts that exploit linguistic and cultural nuance rather than translating English attacks word-for-word.
  • Method over improvisation — can you describe a framework, benchmark, or coverage strategy you've used, and explain why a random-hack approach undertests a model?
  • Judgment on severity and reproducibility — distinguishing a genuine safety failure from a one-off sampling artifact, and knowing when a finding is worth escalating.

Logistics and content exposure

Fully remote and asynchronous; you take batches against deadlines rather than sitting in a shift. Contributors commonly work 10–25 hours per week, though volume fluctuates with customer demand and no minimum is guaranteed. Pay in the $48–62/hr band is observed on comparable Mercor safety projects and varies by specialization and task type.

The work involves reading and generating content touching bias, misinformation, harassment, and other harmful behaviors. All content is text-based. Higher-sensitivity streams are opt-in, topics are disclosed before exposure, and wellness resources are provided — declining a sensitive stream does not disqualify you from the project.