What the work involves
You will be handed AI-generated artifacts — most often Excel or Google Sheets workbooks, sometimes accompanied by supporting documents or slide decks — and asked to judge whether they would survive contact with a real user. That means tracing formulas back to their inputs, checking that ranges did not silently drift, spotting hardcoded values masquerading as calculations, verifying that cross-sheet references resolve, and confirming that named ranges, data validation, and conditional formatting actually behave as described. You then write structured feedback: what is wrong, where, why it matters, and how severe it is against the rubric you have been given.
Tasks arrive in batches and vary in length. Some are quick correctness passes on a single tab; others involve auditing a multi-sheet model or comparing two candidate outputs and defending a preference. Reviewers who do well here treat rubrics as instruments rather than suggestions — they calibrate to the stated scale instead of drifting toward their own house style, and they flag ambiguity in the rubric rather than quietly resolving it.
What the platform screens for
- Verifiable depth. Expect follow-ups on specific workbook failure modes: volatile functions, circular references, array formula spill behavior, INDEX/MATCH versus XLOOKUP tradeoffs, and what breaks when a sheet is copied.
- Written clarity. Feedback is the deliverable. Vague or unstructured commentary is the most common reason strong practitioners fail screening.
- Judgment calibration. Can you distinguish a cosmetic nit from an error that produces a wrong number, and score accordingly?
- Tooling fluency across suites. The listing explicitly calls out Slides and PowerPoint alongside spreadsheets; presentation-layer review is part of the queue.
Logistics
Fully remote, hourly, asynchronous. Observed pay for this listing sits in the $80–120/hr range — Mercor sets rates per project and per reviewer, and nothing is guaranteed. Volume fluctuates with client demand; most reviewers treat this as part-time supplementary work rather than a stable full load. There is usually a paid or unpaid calibration set before live tasks begin, and continued access depends on agreement with gold-standard scores.