What the work involves

You will run adversarial sessions against conversational models and agents, working in both English and Swedish. A typical batch might ask you to attempt a category of jailbreak across ten seeded scenarios, escalate across multiple turns, and record where the model's refusal behavior degrades. Cross-lingual work matters here: attacks that fail in English often succeed in Swedish, and part of the value you provide is documenting that asymmetry — idiom, code-switching, dialect, and translation-layer gaps included.

Every successful probe becomes an artifact. You classify the failure against a provided taxonomy, annotate severity, and write a reproduction path clear enough that an engineer who does not speak Swedish can verify it. Random clever hacks are less useful than structured coverage; reviewers weight consistency with the playbook alongside novelty.

  • Multi-turn manipulation, prompt injection, misuse elicitation, bias exploitation
  • Failure annotation and vulnerability classification against fixed taxonomies
  • Reproducible write-ups: attack case, transcript, severity rationale
  • Flagging systemic patterns rather than one-off outputs

What the platform screens for

Mercor's screen is AI-led and expects verifiable specifics. Expect follow-ups on any red teaming, security, trust-and-safety, or adversarial research you claim — what the scope was, what you found, how you documented it. Native-level Swedish is a hard requirement and is probed conversationally, not self-reported. Candidates from cybersecurity, content moderation, disinformation research, linguistics, psychology, and investigative writing all clear this bar; what they share is the habit of working from a framework and explaining risk plainly to non-specialists.

Logistics

Fully remote and asynchronous, with work delivered in batches. Contributors commonly report 10–25 hours per week, though volume moves with customer demand. Pay observed in this band is $48–62/hr; rates are set per project and are not guaranteed. Higher-sensitivity content streams — harassment, self-harm, extremism, disinformation — are opt-in, disclosed by topic before you see anything, and supported by written guidelines and wellness resources.