What the work actually involves
You write technical prompts in Turkish that probe how a frontier model handles specialized scientific material — synthesis routes, laboratory procedures, pathogen biology, radiochemistry, energetic materials — and then evaluate what comes back. Evaluation is two-sided: you judge scientific accuracy at the level of someone who has run the experiment, and you judge whether the response was appropriate to give at all. A large share of the task is classification: applying a written guideline to decide which category a prompt or conversation falls into, and doing it the same way twice. Because the target language is Turkish, you are also implicitly testing whether the model's safety behaviour holds up outside English — the failure mode this project exists to catch.
What the screen is looking for
- Real PhD depth in one subfield, not survey-level familiarity across many. The listing explicitly prefers demonstrated depth — organic or inorganic synthesis, molecular biology, microbiology, virology, genetics, radiochemistry, nuclear chemistry, radiation physics.
- Genuine Turkish scientific register. Expect to be asked how you handle terminology that Turkish-language labs commonly leave in English, and to write Turkish that reads as though a Turkish researcher wrote it.
- Dual-use judgment. Screeners probe whether you can distinguish an educational question from an operationally useful one, and whether you can explain a refusal decision without either over-blocking legitimate science or hand-waving the risk.
- Guideline discipline. Can you follow someone else's rubric when you personally disagree with it, and flag the disagreement through the right channel rather than quietly grading your own way?
Logistics
Fully remote, asynchronous, roughly 7 hours per week with an immediate start. Basing in Türkiye or the wider MENA region is a stated preference, not a gate — applicants elsewhere are welcome. Pay of $30–34/hr reflects rates observed on this platform for PhD-level bilingual safety work and is not a guarantee; Mercor typically confirms the rate per project. No AI/ML experience is expected — training on the annotation workflow is provided — so the screen weighs domain depth and language far more heavily than tooling familiarity.