What the work involves
You'll receive AI-generated work products — a blameless postmortem for a cascading cache failure, a spreadsheet modeling error budget burn, a slide deck pitching an on-call rotation redesign — and score them against a rubric while writing the reasoning behind each score. The judgment being captured is yours: whether the stated root cause actually follows from the timeline, whether an SLI is measurable from real telemetry, whether remediation items are specific enough to be assigned and closed, whether a severity classification matches the described customer impact. Sessions are usually a mix of scoring, error annotation, and short written critiques that a model can learn from.
Most tasks reward the failure modes only a practitioner notices: a runbook that assumes a dashboard that wouldn't exist mid-outage, a postmortem that quietly blames an individual, an availability target that's mathematically incompatible with the dependency chain described two slides earlier, or a five-nines claim built on a single-region deployment. Presentation quality counts too — malformed tables, unreadable timeline diagrams, and slides that bury the impact statement are all in scope.
What the platform screens for
- Verifiable operational history. Expect follow-ups on specific incidents you ran or wrote up: severity, blast radius, detection path, what the postmortem actually changed.
- Depth under pressure. Interviewers push on SLO math, error budget policy, paging philosophy, and the difference between contributing factors and root cause.
- Rubric discipline. Can you separate "I'd have done it differently" from "this is wrong," and defend a score with cited evidence from the artifact?
- Writing. Feedback needs to be structured, specific, and readable by someone who wasn't in the room.
- Tooling fluency. Real comfort in Google Slides and PowerPoint, plus Sheets/Excel, since a share of the artifacts are decks and models rather than prose.
Logistics
Fully remote and asynchronous, contracted hourly through Mercor. Volume varies by project; many evaluators run 10–20 hours a week around a full-time role, with occasional windows where more work is available. Pay in the $80–120/hr band reflects rates observed on comparable Mercor evaluation projects and is not guaranteed for any given engagement — the offered rate depends on the project, your assessed depth, and screening outcome.