The work

Each task starts with sourcing: you find a publicly accessible Korean PDF in an assigned document type — a newspaper page, a textbook spread, an exam paper, a flyer, form, manual, menu, brochure, notice or worksheet — that contains at least one multimodal element such as a table, diagram, image or handwriting, and you record where it came from. You then build a complete structural map of the page: bound every meaningful region, assign each a component type (document title, section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field), give each a reading-order index, and link dependent regions to their parent figure or table. Finally you transcribe every text region exactly as it appears in Hangul — hanja preserved where it appears, handwriting included, illegible regions flagged rather than guessed — and record page metadata including language, document type, source, dimensions and flags for tables, formulas and handwriting.

The corpus deliberately targets what parsing and vision-language models handle worst, so a typical page is not a clean single-column document. Expect dense multi-column newspaper layouts where reading order crosses columns, exam papers with question and answer-field structure, and handwritten forms. None of the output is machine-generated: component identification, typing, reading order and transcription are all human-authored.

What the screen looks for

Mercor's screening is AI-led and follow-up heavy. Expect it to verify native Korean fluency and full command of Hangul in practice — jamo composition, hanja recognition in older or formal material, Korean input methods — rather than take a claim at face value. It will also probe consistency: whether you can apply a fixed taxonomy the same way across hundreds of pages instead of improvising per document, and how you behave at edges (a caption that could be a paragraph, a smudged handwritten field, a table that spans a page break). Prior work in bilingual transcription, translation, editorial production or AI training data helps; reviewer experience helps more, since experienced annotators are moved onto second-expert review.

Logistics

  • Fully remote and asynchronous, with no fixed shifts described.
  • Open only to contributors based in South Korea, Japan, India, the United States, Canada or Western Europe.
  • Observed rate for this listing is $34/hr — a band reported for this project, not a guarantee.
  • Accuracy is the first measure of quality; handling time is tracked alongside it, not ahead of it.
  • You will need reliable access to public Korean document sources and a working Korean input method.