The work

Each task starts with a publicly available Telugu PDF page in an assigned document type — a newspaper, a textbook spread, an exam paper, a flyer, a form, a menu, a manual. You source the document yourself and record where it came from, then build a complete structural map of one page: bounding boxes around every meaningful region, a component type for each (title, section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field), a reading-order index, and a parent identifier linking captions and cells back to the figure or table they belong to. You then transcribe every text region exactly as it appears, in Telugu script rather than transliteration, including handwritten content, flagging anything illegible instead of guessing.

The corpus is deliberately weighted toward what document models fail on: handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages. Nothing is pre-parsed — component identification, typing, reading order and transcription are all human-authored. Experienced annotators also take on review, checking a colleague's page end to end.

What the screen looks for

  • Native Telugu with full script command. Diacritics, conjunct forms, numeral variants — a single wrong vottu or gunintam counts as a defect. Expect the screen to probe orthography and Unicode handling, not just conversational fluency.
  • Document-handling history. Annotation, transcription, translation, localization, subtitling, proofreading, journalism, OCR post-editing, or regional-language data review. Be ready to name formats and volumes.
  • Taxonomy discipline. The platform wants evidence you apply a fixed label set consistently across hundreds of pages rather than reasoning fresh each time, and that you escalate ambiguous cases instead of inventing categories.
  • Layout judgment. How you determine reading order across newspaper columns, sidebars, boxed callouts and exam answer fields is a real evaluation question here.

Logistics

Remote and asynchronous, open only to contributors based in India, the United States, Canada or Western Europe. Work is task-based rather than shift-based, so hours are largely your own, though accuracy is the first measure tracked and handling time is monitored alongside it. You will need a Telugu input method configured on your machine and the ability to find and download public PDFs. The $13/hr figure is what has been observed on this listing; pay on Mercor projects varies by assessment outcome and is not guaranteed.