What the work actually involves

You work against a model behind a testing interface, usually with a policy document and a target domain assigned for the session. The job is to find the boundary — where a model refuses correctly, where it refuses too much, and where it complies with something it shouldn't. Most attempts fail, which is the point; the deliverable is the small share that succeed, documented well enough that someone else can reproduce the failure and a trainer can build a fix from it.

A typical write-up includes the attack framing, the full transcript, why the output violates policy, how many attempts it took, and whether variations of the prompt still work. Multi-turn escalation, role-play framings, encoding tricks, hypothetical wrappers, and legitimate-professional-context framings are all standard tools. Depth in one grey-area domain matters more than breadth: someone who knows what actually constitutes uplift in biosecurity or what a working exploit chain looks like writes findings that researchers can act on.

What the screen looks for

  • Verifiable background. Five-plus years in AI safety, trust & safety, cybersecurity, investigative journalism, or a life science — with specifics about what you handled and at what level of severity.
  • Domain depth under follow-up. Expect the interview to push past your first answer on your claimed specialty; general familiarity reads as thin quickly.
  • Calibration. Whether you can distinguish a genuinely harmful output from something merely uncomfortable, and whether you recognize over-refusal as a failure mode too.
  • Written clarity. Red teaming output is a document. Vague or dramatized reports have little value.

Logistics

Remote and asynchronous, contract-based, with hours logged against assigned campaigns. Most contributors run 10–25 hours weekly; volume fluctuates as campaigns open and close. Work is confidential, and NDA plus attestation to handling sensitive material is standard. Content is deliberately unpleasant in places — misinformation, fraud, self-harm, CBRN adjacency — and the platform expects you to be honest with yourself about that before applying.