What the work involves

You spend your sessions adversarially probing conversational models and agents, then writing up what broke. A typical task starts from a taxonomy or playbook — a category of harm, a misuse pattern, an attack class — and asks you to construct prompts in English or Assamese that push the model past its guardrails. When a model fails, you annotate the failure, classify the vulnerability, and produce a reproducible case: the exact prompts, the turn sequence, the conditions under which it recurs. Multi-turn manipulation matters here; single-shot jailbreaks are the easy part, and much of the value comes from attacks that build slowly across a conversation or exploit the gap between a model's English and Assamese safety behavior.

The Assamese requirement is substantive, not cosmetic. Safety training is unevenly distributed across languages, and low-resource-language red teaming surfaces failures that English testing never reaches — code-switching, transliteration, register shifts, culturally specific framings of harm. You need native fluency in both languages to know when an output is subtly wrong rather than obviously wrong.

What the platform screens for

Mercor's screening is AI-led and interview-based, and it presses on specifics. Expect to be asked to describe an actual attack you constructed and why it worked at a mechanistic level, not just that it did. Vague claims of "red teaming experience" don't survive follow-up questions. Reviewers also look for structured thinking — whether you work from frameworks, taxonomies, or benchmarks rather than improvising — and for writing quality, since your deliverable is documentation another person has to act on. Verified fluency in Assamese is confirmed during the process.

Logistics and content exposure

  • Fully remote and asynchronous; you claim and complete tasks on your own schedule.
  • Observed pay band is $20–22/hr, stated as reported by contributors — rates vary by project and are not guaranteed.
  • Hours are typically flexible and project-dependent, with volume rising and falling as customer engagements start and end.
  • The work involves reviewing outputs touching bias, misinformation, and harmful behaviors. Topics are disclosed before exposure, higher-sensitivity projects are opt-in, and wellness resources are provided.