What the work actually involves

You write expert-level prompts in Japanese across your scientific specialty — the kind of question a senior researcher would ask a colleague, not a textbook exercise — and then judge how the model answers. A single task usually means drafting a technically precise prompt, reading a model response closely, and annotating it against a rubric that covers scientific accuracy, usefulness to a knowledgeable reader, and whether the model handled sensitive or dual-use material appropriately. You will also apply structured classification guidelines to prompts and conversations, which means reading someone else's taxonomy carefully and applying it the same way twice.

The safety dimension is central, not decorative. Much of the material sits in chemical, biological, radiological and nuclear territory, so you are constantly deciding where the line falls between information a competent graduate student would find in a standard reference and information that meaningfully lowers a barrier to harm. The interesting cases are rarely obvious — a model that refuses a routine question about reaction stoichiometry is failing just as clearly as one that volunteers a synthesis route it shouldn't. Working in Japanese adds a layer: technical terminology, katakana loanwords, and register all affect whether a prompt is genuinely testing the model or just confusing it.

What the platform screens for

Mercor's screen is AI-led and conversational. Expect it to establish your PhD field and stage, then push on subfield depth — specific techniques, instruments, failure modes in your own experimental work — because breadth claims collapse quickly under follow-up. It will test Japanese scientific fluency directly, not just conversational ability, and check that your written English is good enough to follow English-language annotation guidelines. Prior AI/ML experience genuinely isn't required; judgment about dual-use information and the discipline to apply a rubric consistently are what carry weight.

Logistics

  • Fully remote and asynchronous; work is claimed from a queue rather than scheduled in shifts.
  • Start date is immediate, with training on the annotation workflow provided.
  • Contributors commonly commit 10–20 hours per week, though volume varies with project demand.
  • Being based in Japan or East Asia is a stated preference, not a gate — applicants elsewhere are explicitly welcome.
  • Observed pay band is $68–72/hr; rates on Mercor vary by project and are not guaranteed.