What the work actually involves

You write expert-level prompts in Danish inside your own scientific territory — synthesis routes, assay design, pathogen handling, radiochemistry, energetic materials — and then judge what the model gives back. Judging means three things at once: is the chemistry or biology correct, is the answer actually useful to a competent scientist, and does the model draw the right line when the question edges toward dual-use territory. A third stream of work is classification: applying a written taxonomy to prompts and conversations so that labels stay consistent across dozens of annotators who never speak to each other.

The Danish requirement is load-bearing, not cosmetic. Much of the value comes from finding places where a model that is careful and competent in English becomes sloppy, over-cautious, or subtly wrong once the same technical content arrives in Danish — including where it mishandles Danish scientific register, borrows English terminology awkwardly, or fails to recognise a hazardous request that was obvious in English.

What the screen is looking for

  • Verified doctoral depth in one subfield, not a survey-level tour of chemistry and biology. Screeners follow up hard on whatever you claim, and a named technique with named failure modes travels much further than a broad topic list.
  • Genuine Danish scientific fluency — the ability to write technically precise Danish prompts, not conversational Danish plus English jargon.
  • Calibrated dual-use judgment. The platform wants people who can distinguish a legitimate methods question from a request for operational uplift, and who can articulate why, rather than people who refuse everything or nothing.
  • Guideline discipline — willingness to follow a rubric you disagree with while flagging the disagreement through the right channel.

No prior AI or ML background is expected; the workflow is taught. Prior experience grading, peer-reviewing, or red-teaming technical content is a plus rather than a gate.

Logistics

Fully remote and asynchronous, roughly 7 hours per week, start immediate. Denmark or wider Western Europe is a stated preference only — applicants elsewhere are explicitly welcome. Pay of $61–65/hr reflects rates observed on this listing; actual offers vary by assessment outcome, task type, and project phase, and are not guaranteed. Work is typically delivered in batches with quality review, so throughput matters less than consistency across a batch.