The work
You'll receive AI-generated artifacts — a KPI dashboard mockup, a monthly performance report, a spreadsheet model behind a scorecard, or an executive deck summarizing business metrics — and grade them against a structured rubric. Most tasks ask you to do three things: judge whether the analysis is defensible, mark specific errors with location and severity, and write a short justification another reviewer could follow without your context.
The errors that matter here are often quiet ones. A chart type that misrepresents a trend, a metric definition that silently changes between slides, a percentage-of-total that doesn't reconcile, a filter that drops nulls without disclosure, a dashboard that buries the one number an executive actually needs. You're expected to catch those, not just typos — though presentation and formatting defects are in scope too, since these outputs are graded on whether a stakeholder could use them as-is.
What the screening looks for
Mercor's screen is AI-led and follow-up heavy. It probes for verifiable specifics: which BI tools you built in, what metrics you owned, who consumed your reporting, and what you changed when a dashboard wasn't being used. Expect to be asked to reason aloud about a flawed report rather than describe your philosophy of good design. Vague seniority claims tend to collapse under the second or third follow-up; concrete examples with numbers and constraints hold up.
Logistics
- Fully remote, asynchronous, hourly.
- Task volume varies by project cycle; many evaluators work 10–20 hours per week, some more during ramp-ups.
- No guaranteed hours or continuity — engagements are project-based and rates are as observed, not promised.
- You'll work in Google Slides, Sheets, PowerPoint, and Excel constantly, so genuine fluency there is a hard requirement, not a resume line.