The work

This is adversarial testing of conversational AI, done bilingually. You'll run structured attack sessions against models and agents — single-turn jailbreaks, prompt injection through tool or document context, misuse framings, bias elicitation, and slow multi-turn manipulation where the model concedes ground gradually. A large part of the value here is cross-lingual: safety training frequently generalizes poorly from English to Vietnamese, and attacks that fail in one language often succeed when translated, code-switched, or reframed through Vietnamese cultural and political context. Native fluency in both languages is a hard requirement, not a preference.

Every finding becomes an artifact. You annotate the failure, classify it against the project taxonomy, note reproduction steps and the model's behavior across retries, and write it up so a non-specialist can understand why it matters. Random one-off hacks are less useful here than coverage against a benchmark — the platform is explicit that structured probing beats improvisation.

What the screen looks for

  • Verifiable prior red teaming. AI adversarial work, penetration testing, exploit development, trust-and-safety abuse analysis, or socio-technical probing. Expect follow-ups on specific attacks you've run and what the system did.
  • Genuine bilingual depth. Screens for Vietnamese language roles typically test register, idiom, and regional variation — not just conversational ability. Translation-tool fluency does not pass.
  • Judgment under ambiguity. Whether a borderline output counts as a failure, how you'd rate severity, and when a refusal is over-cautious rather than safe.
  • Sensitive-content readiness. The work touches bias, misinformation, self-harm, and harassment material. Topics are disclosed before exposure and higher-sensitivity projects are opt-in, but you should be honest with yourself about fit.

Logistics

Fully remote and largely asynchronous, with work assigned in project batches that rotate across customers. Observed pay for this listing runs $17–25/hr, with placement in the band typically reflecting prior security or adversarial-ML background. Hours are flexible and self-scheduled; contributors commonly report 10–25 hours per week, though volume depends on active customer projects rather than a fixed commitment. Expect calibration materials, a taxonomy document, and periodic guideline updates you're responsible for keeping current with.