What the work involves

You spend your sessions trying to break a model rather than use it. That means constructing adversarial prompts in both English and Thai, chaining multi-turn conversations that erode a model's refusals, testing prompt-injection paths through tool use or retrieved content, and probing where a safety policy holds in one language but collapses in the other. Cross-lingual gaps are a large part of why this role is bilingual: a guardrail trained mostly on English data often behaves differently when the same request arrives in Thai, in transliterated Thai, or in a code-switched mix.

When an attack lands, the artifact matters more than the exploit. You annotate the failure, classify it against the project's vulnerability taxonomy, judge severity, and write it up so an engineer who has never seen your conversation can reproduce it. Consistency is enforced through playbooks and benchmarks — the goal is a dataset customers can act on, not a collection of one-off clever tricks.

What the screen looks for

  • Verifiable red teaming history — adversarial ML work, penetration testing, trust-and-safety abuse analysis, or socio-technical probing, described with specifics rather than tool names.
  • Genuine native-level fluency in both languages, including register, slang, and the ways Thai speakers actually phrase evasive or coded requests.
  • Structured thinking under follow-up. Expect the interviewer to push on how you'd systematically cover an attack surface, not just what your best jailbreak was.
  • Write-up quality. Reproducibility, clear severity reasoning, and plain-language risk explanation carry real weight.

Logistics

Fully remote and asynchronous, contractor terms, with contributors typically choosing their own hours against weekly volume commitments. Work is entirely text-based. Higher-sensitivity workstreams — bias, misinformation, harmful behaviors — are opt-in, with topics disclosed before exposure and wellness resources provided. Project availability shifts with customer demand, so hours can vary week to week; the $24–35/hr band reflects rates observed on this listing and is not a guarantee.