The work

You attack models on purpose. A typical session means picking a target behavior — a policy the model is supposed to hold, a refusal it should maintain, a factual claim it should not fabricate — and then working through structured attempts to break it: role-play framing, prompt injection through pasted content, incremental multi-turn escalation, translation and code-switching between English and Finnish, and exploitation of bias or cultural blind spots. When something breaks, you write it up so someone else can reproduce it: the exact prompts, the turn sequence, the classification against the project taxonomy, and an assessment of severity.

The Finnish requirement is load-bearing. A great deal of safety tuning is done in English, and models frequently behave differently once a request crosses into a lower-resource language. You will be probing exactly that gap — testing whether Finnish phrasing, idiom, morphology, or culturally specific framing routes around guardrails that hold in English. Native fluency in both languages is what makes the comparison meaningful.

What the screen looks for

  • Prior adversarial work you can describe concretely. AI red teaming, penetration testing, exploit development, trust-and-safety abuse analysis, or disinformation research. Expect follow-ups asking what you actually tried and what failed.
  • Method over improvisation. Interesting one-off jailbreaks matter less than being able to explain your coverage strategy and why you tested what you tested.
  • Documentation discipline. Reproducibility is the deliverable. Vague findings have no downstream value.
  • Genuine bilingual depth, including register, slang, and regional variation — not textbook Finnish.

Logistics and content

Remote and asynchronous, with contributors typically committing 10–25 hours per week against project deadlines rather than fixed shifts. Everything is text-based. The material touches bias, misinformation, harassment, and harmful behaviors; higher-sensitivity tracks are opt-in, topics are disclosed before exposure, and wellness resources are provided. The stated band is $48–62/hr as observed on the platform — actual rates vary by specialty, project, and assessment outcome, and are not guaranteed.