What the work involves
You spend most of your time writing prompts in Japanese — not casual questions, but carefully constructed ones that probe how a model behaves around sensitive and dual-use subject matter. Then you work in the other direction: reading prompts and multi-turn conversations against a written guideline set and assigning classifications, flagging adversarial phrasings, indirect requests, roleplay framings, and escalation across turns. Every judgment needs a short written rationale in English that another reviewer could follow without asking you what you meant.
Much of the value you add is cultural rather than purely linguistic. A request that reads as innocuous in a literal English gloss may carry a very different weight in Japanese — through keigo and register shifts, indirection, euphemism, slang from Japanese-language forums, or references that only make sense inside a Japanese legal, medical, or social context. Models trained largely on English data miss these, and the point of the role is to surface them.
What the platform screens for
Mercor's screen is AI-led and conversational. Expect it to test Japanese fluency directly, ask you to reason aloud about a borderline case, and probe whether you can apply a rule as written rather than substituting your own instincts. It will also check that your English writing is clear enough to carry a rationale, since documentation is the deliverable. Consistency matters more than cleverness — annotation work is judged on whether two reviewers reading your notes would land in the same place.
Logistics
- Remote and asynchronous; work is completed on your own schedule against deadlines.
- Roughly 7 hours per week, part-time, with an immediate start.
- Observed band for this listing is $48–52/hr; pay is as reported, not guaranteed, and can vary by project and assessed level.
- Based in Japan or the wider East Asia region is a stated preference only — applicants elsewhere are welcome.
- Bachelor's degree completed or in progress.