What the work actually involves

You write prompts in Chinese at the depth a working researcher would ask — a synthesis question with a real reagent constraint, a cloning protocol that fails for a specific reason, a question about detector calibration or shielding geometry. Then you read what the model returns and judge it on three axes at once: is the science correct, is it actually useful to the person asking, and does it handle sensitive or dual-use material the way the guidelines say it should. A third strand of the work is classification — applying a written rubric to sort prompts and whole conversations into categories, consistently, across hundreds of items.

The hard part is rarely the chemistry. It is holding a rubric steady across a long queue, and distinguishing a response that is dangerous from one that merely sounds alarming — textbook hazard information, published protocols, and genuine uplift are different things, and conflating them makes an annotator less useful, not safer.

What the screen looks for

  • Verifiable doctoral depth in one subfield. Expect follow-ups that go two or three layers past your stated specialty — mechanism, failure modes, what you would actually run in the lab. Broad coverage reads weaker than narrow depth here.
  • Genuine Chinese scientific register. Not conversational fluency: the ability to write technical prompts using the terminology a Chinese-language researcher or textbook would use, including where translation is contested.
  • Calibrated dual-use judgment. The screen will probe whether you can articulate why something crosses a line, not just that it feels wrong.
  • CBRN breadth from a chemistry base is a plus — radiochemistry, nuclear chemistry, energetic materials, radiation physics all count toward the R/N side.

Logistics

Fully remote and largely asynchronous, with work drawn from a queue rather than scheduled shifts. Start date is immediate; contributors commonly commit 10–20+ hours weekly, and throughput tends to matter more than fixed hours. Hong Kong or the wider East Asia region is a stated preference but explicitly not a requirement. Onboarding includes training on the annotation workflow, so the AI/ML side is learned on the job; the domain expertise and the Chinese are what you bring. Pay is observed at $68–72/hr and is not guaranteed — bands on Mercor vary by project, assessed depth, and language pair.