The work

You receive AI-produced deliverables that imitate real operations work: a spreadsheet computing safety stock and reorder points, a capacity model for a bottleneck work center, a network study comparing DC footprints, an S&OP review deck for an executive audience. Your job is to decide whether the artifact would survive contact with an actual planning team. That means checking formulas and units, testing whether the stated assumptions actually support the conclusion, catching invented benchmarks or misapplied EOQ and Little's Law reasoning, and flagging slides where the chart contradicts the headline.

Feedback is written, structured, and scored against a rubric. Volume-heavy commentary is not the goal — evaluators who succeed here write tight, specific critiques that point to the exact cell, line, or claim at fault and explain the operational consequence of the error.

What the screening measures

  • Verifiable depth. Expect follow-ups on how you actually set service levels, sized buffers, or resolved a capacity constraint in a named role, not textbook definitions.
  • Error-finding instinct. You may be shown a flawed model or deck excerpt and asked what is wrong and what matters most.
  • Judgment calibration. Can you separate a cosmetic formatting issue from an assumption error that invalidates the output?
  • Written clarity. Short-answer prompts are read for precision and structure, since the deliverable is prose feedback.
  • Tooling fluency. Real comfort in Excel/Sheets and PowerPoint/Google Slides is checked, not assumed.

Logistics

Fully remote and asynchronous. Task batches are claimed as they appear; most evaluators work in blocks of a few hours rather than fixed shifts, and availability of 10–20 hours per week is typical. Pay is hourly at the band observed for this listing — not guaranteed, and subject to task type and platform calibration.