The work
You receive AI-generated artifacts — a research memo on Baroque patronage networks, a slide deck pitching a museum exhibition concept, a spreadsheet cataloguing archival holdings, a comparative essay on translation theory — and you grade them. Grading means more than a score: you mark where the model invented a citation, flattened a contested interpretation into false consensus, misattributed a work, mangled a transliteration, or produced a deck that is factually sound but visually incoherent. Written feedback is the deliverable, and it needs to be structured enough that a non-expert reviewer can follow your reasoning to the specific line or slide.
Much of the difficulty is calibration rather than knowledge. Humanities outputs rarely fail in the way a math proof fails; they fail by being plausible, fluent, and subtly wrong about provenance, period, register, or scholarly consensus. Rubrics ask you to separate "an interpretation I disagree with" from "an interpretation no serious scholar would defend," and to hold that line consistently across dozens of samples.
What the screen looks for
- Verifiable professional history: five or more years in art history, literature, classics, musicology, curation, criticism, cultural heritage, philosophy, or an adjacent field. Advanced degrees are preferred, not required.
- Real depth under follow-up — the AI interviewer will push past your first answer to see whether you can cite specifics, name sources, and explain where the field disagrees with itself.
- Genuine fluency with Slides and PowerPoint. Deck evaluation is a large share of the work, and candidates who treat presentation software as an afterthought tend not to pass.
- Evaluative temperament: the ability to give a low score with a clear reason and a high score without hedging.
Logistics
Remote, hourly, asynchronous. Work arrives in batches and volume fluctuates with client demand — some weeks offer 20+ hours, others little. Most contributors set their own schedule and are asked to commit to a minimum weekly availability rather than fixed hours. Pay bands of $80–120/hr have been observed for this domain on Mercor; the actual rate is set per engagement and is not guaranteed.