What the work involves
You spend your sessions trying to break a model on purpose. That means constructing adversarial prompts in both English and Bengali, running multi-turn manipulation attempts, testing prompt injection through tool or document context, and probing where safety training thins out under translation, code-switching, or transliterated Bengali script. When an attack lands, the deliverable is not the screenshot — it is a reproducible artifact: the exact prompt chain, the failure classification against the project taxonomy, a severity judgment, and a note on why the model failed rather than just that it did.
Much of the day-to-day is annotation and classification discipline. You will work inside taxonomies and playbooks rather than freestyling, tag failure modes consistently across sessions, and flag patterns that look systemic rather than one-off. Some tasks touch sensitive content — bias, misinformation, harassment vectors, harmful instruction requests. Mercor states that higher-sensitivity queues are opt-in, topics are disclosed before exposure, and wellness resources are available.
What the screen looks for
- Genuine bilingual depth. Native-level English and Bengali, including register shifts, regional variation, and Romanized Bengali — not translation-app fluency.
- Prior adversarial work. AI red teaming, penetration testing, exploit development, trust-and-safety abuse analysis, or socio-technical probing. Expect follow-ups asking you to walk through a specific attack you ran and why it worked.
- Structured method over clever one-offs. Interviewers probe whether you can generalize an attack into a test class and describe coverage.
- Clear written reasoning. Your report is the product; ambiguity in your write-up is a defect.
Logistics
Fully remote and asynchronous, contractor engagement, work claimed from queues rather than assigned on a fixed schedule. Volume fluctuates by customer project, so treat it as variable-hour rather than steady full-time. Observed pay for this listing is $20–22/hr; rates are set per project and are not guaranteed. Expect a calibration period with sample tasks and reviewer feedback before full-rate work opens up.