What the work involves
You receive AI-generated artifacts — a pitch deck, a board memo, a financial model, a formatted report — and evaluate them the way a demanding production or design QA reviewer would before anything reaches a client. That means checking claims against source material, catching broken cross-references and inconsistent number formatting, spotting misaligned objects, mismatched type hierarchies, off-brand color use, orphaned text boxes, and slides that technically answer the prompt but would never survive a partner review. Each task ends with a rubric score plus written feedback specific enough that a model trainer can see exactly what failed and why.
Most projects mix scoring passes with side-by-side comparisons of two candidate outputs, and occasionally rewriting or rebuilding a deck to demonstrate the correct standard. Volume and task type shift as projects rotate.
What the platform screens for
- Verifiable production history — the environments you worked in (consulting, agency, in-house comms, investor relations, publishing), the deck volume and turnaround pressure you handled, and who reviewed your work.
- Tool depth under follow-up — master slides and layouts, theme inheritance, PowerPoint vs. Google Slides behavior differences, table and chart formatting, exported PDF fidelity, spreadsheet formula and formatting hygiene.
- Evaluation judgment — whether you can separate a taste preference from a defect, rank multiple flaws by severity, and articulate a standard rather than a reaction.
- Writing quality — feedback that is concrete, ordered, and free of vague adjectives.
Logistics
Fully remote, asynchronous, contractor engagement billed hourly. Contributors commonly commit 10–20 hours per week, self-scheduled, though some projects offer more during ramp periods. Expect a short paid calibration phase where your scores are compared against reviewer consensus; ongoing work depends on staying inside that band. Reliable access to both Microsoft Office and Google Workspace is required.