What the work actually involves

You spend most of your time in two modes. The first is prompt authoring: writing questions in Czech that only someone with graduate training in your subfield could write — the kind that require real reagent knowledge, plausible protocol detail, or awareness of where a synthesis or assay actually breaks. The second is evaluation: reading model responses and judging whether they are scientifically correct, genuinely useful to a competent researcher, and appropriately handled where the topic touches dual-use territory. A third strand is classification — applying a written taxonomy to label prompts and conversations consistently, including edge cases the guidelines did not anticipate.

Because the subject areas span chemical, biological, and radiological/nuclear ground, the safety judgment is not decorative. You will regularly meet prompts where the correct answer is neither "refuse" nor "answer fully," and the annotation asks you to say why. No AI or ML background is required; the platform trains the workflow. What cannot be trained in a week is the doctoral-level domain instinct and the Czech.

What the screen looks for

  • Verifiable doctoral depth in one subfield, not a survey-level tour. Expect follow-ups that push past your first answer into mechanism, instrumentation, or failure modes.
  • Genuine Czech scientific register. Czech technical writing is not translated English; screeners look for whether you can pose a graduate-level chemistry or biology question in Czech without reaching for calques, and whether you can flag when a model's Czech is fluent but technically wrong.
  • Calibrated dual-use judgment — the ability to distinguish a hazardous-detail request from a legitimate one, and to articulate the line rather than gesture at it.
  • Rubric discipline — willingness to follow a guideline you disagree with while logging the disagreement.

Logistics

Fully remote and asynchronous, with an immediate start and a part-time commitment of roughly 7 hours per week. Observed pay for this listing sits at $53–57/hr; rates on Mercor vary by project and are not guaranteed. Being based in the Czech Republic or wider Eastern Europe is a stated preference only — applicants elsewhere are explicitly welcomed. Work is self-scheduled around annotation batches, so consistency across the week matters more than fixed hours.