What the work involves

You receive an AI-generated artifact — a spreadsheet with a TAM/SAM build, a competitive landscape deck, a win/loss synthesis, a survey crosstab summary — along with the prompt or brief it was produced from. Your job is to judge whether it would survive contact with a real client or internal stakeholder. That means checking whether the market sizing logic is sound and its assumptions traceable, whether competitor claims are attributed to plausible sources rather than invented, whether segmentation is coherent, and whether the deck actually communicates a point of view instead of listing facts.

Evaluations are rubric-driven. You score dimensions such as factual accuracy, analytical rigor, source credibility, structure, and presentation quality, then write structured comments explaining each deduction with enough specificity that an engineer or a model can act on it. Slide-level craft matters here more than in most evaluation work: misaligned elements, unreadable charts, inconsistent labeling, and 2x2s that don't hold up are all in scope.

What the platform screens for

  • Verifiable depth: which industries you've covered, which data sources you've actually worked in (Euromonitor, IBISWorld, Gartner, Crunchbase, PitchBook, panel providers), and what kinds of deliverables you owned end to end.
  • Judgment under follow-up: expect probes on how you'd catch a fabricated market figure or a competitor claim that sounds right but isn't.
  • Writing quality: feedback that is specific and calibrated, not vague praise or blanket criticism.
  • Real proficiency in Slides and PowerPoint — this is tested, not assumed.

Logistics

Fully remote and asynchronous. Tasks are claimed from a queue; most contributors work in blocks of a few hours rather than fixed shifts, and volume fluctuates with project cycles. Expect a calibration period with sample tasks and reviewer feedback before full task access. Advanced degrees help but do not substitute for hands-on delivery experience.