What the work actually involves
You work against a model behind a testing interface, usually with a policy document and a target domain assigned for the session. The job is to find the boundary — where a model refuses correctly, where it refuses too much, and where it complies with something it shouldn't. Most attempts fail, which is the point; the deliverable is the small share that succeed, documented well enough that someone else can reproduce the failure and a trainer can build a fix from it.
A typical write-up includes the attack framing, the full transcript, why the output violates policy, how many attempts it took, and whether variations of the prompt still work. Multi-turn escalation, role-play framings, encoding tricks, hypothetical wrappers, and legitimate-professional-context framings are all standard tools. Depth in one grey-area domain matters more than breadth: someone who knows what actually constitutes uplift in biosecurity or what a working exploit chain looks like writes findings that researchers can act on.
What the screen looks for
- Verifiable background. Five-plus years in AI safety, trust & safety, cybersecurity, investigative journalism, or a life science — with specifics about what you handled and at what level of severity.
- Domain depth under follow-up. Expect the interview to push past your first answer on your claimed specialty; general familiarity reads as thin quickly.
- Calibration. Whether you can distinguish a genuinely harmful output from something merely uncomfortable, and whether you recognize over-refusal as a failure mode too.
- Written clarity. Red teaming output is a document. Vague or dramatized reports have little value.
Logistics
Remote and asynchronous, contract-based, with hours logged against assigned campaigns. Most contributors run 10–25 hours weekly; volume fluctuates as campaigns open and close. Work is confidential, and NDA plus attestation to handling sensitive material is standard. Content is deliberately unpleasant in places — misinformation, fraud, self-harm, CBRN adjacency — and the platform expects you to be honest with yourself about that before applying.