What the work involves
You write prompts in Russian that test how a model behaves at the edges: adversarial phrasings, indirect requests, culturally specific framings that an English-first safety policy may not anticipate. You then apply a written taxonomy to classify prompts and conversations — harm category, severity, whether the model's response crossed a line — and write a short justification in English for each call. The justification matters as much as the label; downstream policy work depends on being able to reconstruct why a borderline item landed where it did.
Expect the guidelines to change under you. Taxonomies get revised mid-project, edge cases you flagged in week two become their own category in week four, and reviewers will push back on specific labels. Comfort with re-reading a spec and adjusting rather than defending your first instinct is a real part of the job.
What the screening looks for
Mercor's screen is AI-led and largely conversational. It will test Russian fluency directly — including register, regional variation, and whether you can produce natural adversarial phrasing rather than translated English. It probes judgment on dual-use material: can you distinguish a legitimate question from one engineered to extract harmful specifics, and can you articulate the difference without hedging into uselessness? It also verifies the plain facts: degree status, written English, weekly availability.
Logistics
- Remote, asynchronous, no fixed hours; work is claimed from a queue
- Roughly 7 hours per week, part-time, immediate start
- $38–42/hr observed on this listing — bands on Mercor vary by project and are not guaranteed
- Eastern Europe or other Russian-speaking regions preferred but explicitly not required