The work
You are the adversary. Day to day, that means opening a conversation with a model and methodically trying to make it fail — direct jailbreak attempts, indirect prompt injection, role-play framing, incremental multi-turn escalation, and language-switching attacks that exploit weaker safety coverage in Indonesian than in English. Cross-lingual probing is a large part of why this cohort exists: safety training that holds in English often degrades in Bahasa Indonesia, in code-switched prompts, or in regionally specific slang, idiom, and cultural framing. You will be asked to find exactly those seams.
Every successful attack becomes an artifact. You annotate what failed and why, classify the vulnerability against a provided taxonomy, and write the case up so an engineer who has never seen your conversation can reproduce it. Consistency matters more than cleverness here — the project runs on shared benchmarks and playbooks, and an attack that cannot be replicated is not deliverable data. Expect regular guideline updates, calibration rounds, and reviewer feedback on your write-ups.
What the screen looks for
Mercor's process is largely AI-led: a resume parse, then a structured voice or chat interview, then task-based calibration. The interview probes for verifiable red teaming history — adversarial ML work, penetration testing, trust and safety, abuse analysis, or socio-technical research — and pushes on specifics. Vague claims about "testing chatbots" collapse under follow-up; naming a taxonomy you used, an attack class you specialize in, or a failure you documented does not. Genuine native-level fluency in both English and Indonesian is verified in conversation, not self-reported.
Logistics and content exposure
- Fully remote and asynchronous; you claim tasks against open batches rather than working fixed shifts.
- Hours are variable and project-driven — some contributors work 10–15 hours weekly, others more when a customer engagement ramps.
- Observed pay is $17–25/hr, positioned by experience and language coverage. Rates are as reported, not guaranteed.
- The work involves reading and generating text touching bias, misinformation, and harmful behaviors. Topics are disclosed before exposure, higher-sensitivity streams are opt-in, and wellness resources are provided.