The work
You get an image and a caption written by a model. Your job is to read the caption clause by clause and decide whether the image actually supports each claim. Most defects fall into predictable families: the wrong object named, a count that is off by one or two, a colour that isn't there, text inside the image (signage, packaging, screenshots, handwriting) transcribed incorrectly, and — the most common and most argued-over — claims the image simply doesn't establish, like intent, location, relationship, or time of day. Each flag you file needs to be tied to something visible. "The caption says three cyclists; there are four, one partly behind the van" is a finding. "Feels wrong" is not.
- Work a fixed queue rather than an open-ended stream
- Log every substantiated defect on an item, not just the first one
- Distinguish a false claim from an unsupported one, and from a caption that is merely thin
- Expect some images that are genuinely ambiguous, low-resolution, or crowded
What the screen looks for
Mercor's screening is AI-led and conversational. It will probe whether you can articulate why a caption is defective in terms another reviewer could verify, whether you understand the difference between a caption being incomplete and a caption being wrong, and whether you'll hold a call rather than flag on vibes. Native-level English matters because much of the judgment is linguistic: whether "a man appears to be waiting" is hedged enough to survive an ambiguous image. No specialist degree is required or asked for.
Logistics
Fully remote and asynchronous. Observed pay for this listing is $30/hr; rates on Mercor vary by project and are set at offer, not guaranteed by the posting. Hours are typically flexible within a project window, but queues have deadlines, so expect to commit to a rough weekly volume. Sessions tend to be long enough that eye fatigue is a real factor — reviewers who pace themselves catch more than reviewers who sprint.