What the work actually looks like

You write prompts in Finnish that a competent specialist in your subfield would find genuinely hard — not trivia, but the kind of question where a plausible-sounding wrong answer is easy to produce. Then you evaluate what the model returns: is the chemistry correct, is the reasoning sound, does it hedge where hedging is warranted, and does it handle dual-use material responsibly rather than either over-refusing or answering too helpfully? A third strand of the work is classification: applying a written taxonomy to prompts and conversations, consistently, across many items. Expect the guidelines to be revised mid-project and expect your earlier calls to be re-examined against the new version.

What the screen is looking for

Mercor's screening is AI-led and follow-up driven. It will push on your actual subfield — the techniques you ran, the instruments you used, what failure modes you learned to recognise — because prompt quality tracks depth, not breadth. It also tests Finnish specifically as a technical language: whether you can write precise scientific Finnish rather than translating English phrasing, and whether you know how Finnish-language chemistry and biology terminology is handled in practice (loan terms, established Finnish equivalents, register in academic versus clinical writing). Finally, it probes safety judgment: can you articulate where the line sits between explaining a mechanism and providing an operational route, and can you defend that line under pressure without collapsing into blanket refusal.

Logistics

  • Remote and asynchronous; no fixed hours, work is pulled from a queue.
  • Roughly 7 hours per week, part-time, starting immediately.
  • Based in Finland or Western Europe is a stated preference only — applicants elsewhere are explicitly welcome.
  • No AI/ML experience required; workflow training is provided.
  • $61–65/hr is the observed band for this listing, not a guarantee — rates on Mercor vary by track, project phase, and assessed depth.

The realistic ceiling on this work is guideline adherence. Strong scientists who grade by instinct rather than by rubric tend to get low agreement scores and short engagements; the people who last treat the rubric as the object of study alongside the science.