What the work involves
You will spend most of your time in adversarial conversation with AI models — writing attack prompts, escalating across multiple turns, and looking for the point where safety behavior degrades. Work is split between English and Odia, and a substantial part of the value here is testing whether guardrails that hold in English hold equally in Odia: transliteration tricks, code-switching, culturally specific framings, script-level manipulation, and low-resource-language blind spots.
When you find a failure, the artifact matters as much as the attack. You will annotate what broke, classify the vulnerability against the project taxonomy, note reproducibility conditions, and write it up so a customer engineering team can act on it. Consistency is enforced through playbooks and benchmarks — this is structured probing, not freeform hacking.
What the platform screens for
- Genuine native fluency in both languages, including register, idiom, and regional variation in Odia — expect to be tested, not just asked.
- Prior adversarial experience of some kind: AI red teaming, penetration testing, trust and safety, disinformation research, or abuse analysis.
- Systematic thinking. Screeners probe whether you work from a threat model and coverage plan or just improvise.
- Write-up quality. Vague findings fail; reproducible, well-scoped reports pass.
Logistics and content
Remote, asynchronous, and text-only. Contributors typically set their own hours against weekly throughput expectations; steady part-time availability is generally more useful to the project than sporadic bursts. The work involves reviewing model outputs touching bias, misinformation, and harmful behavior. Mercor states that higher-sensitivity streams are opt-in, topics are disclosed before exposure, and wellness resources are provided. Pay is stated as observed on the platform at $20–22/hr and is not guaranteed for any individual contract.