What the work actually involves
Most of your time goes into two activities. The first is authoring prompts in German across sensitive subject areas — health, self-harm, extremism, weapons, fraud, privacy, and other dual-use territory — that are realistic enough to probe how a model behaves rather than merely tripping a keyword filter. The second is classification: reading prompts and multi-turn conversations, applying a structured guideline document, and writing a short justification for each label. The written reasoning matters as much as the label; a rating with no defensible rationale is unusable to the team downstream.
A meaningful part of the job is noticing what English-language safety tooling misses. German-specific evasion — compound-word obfuscation, dialect and Austrian/Swiss variants, formal register used to launder a harmful request, historically loaded phrasing, code words circulating in German-language online communities — is exactly what you're hired to see. You'll also be asked to flag escalation patterns, where a conversation begins innocuously and drifts turn by turn toward something the model should refuse.
What the screen looks for
Mercor's process is an AI-led interview followed by task-based calibration. The screen tests verifiable facts (fluency and how you acquired it, degree status, region), then pushes on judgment: can you draw a line between information that is uncomfortable and information that is genuinely dangerous, and can you say where the line is and why? Expect follow-ups that reverse your position to see whether your reasoning holds. Candidates who default to "refuse everything sensitive" screen out as readily as those who wave everything through.
Logistics
- Fully remote and asynchronous; no fixed shifts, work delivered against weekly volume.
- Roughly 7 hours per week, part-time, immediate start.
- Germany or wider Western Europe is preferred but explicitly not required.
- Pay observed in the $48–52/hr band; rates on Mercor projects vary by project and are not guaranteed.
- The material is deliberately unpleasant at times — self-harm, violence, exploitation content appears in the queue.