What the work involves
You'll receive AI-generated artifacts that mimic real product launch deliverables: experiment design docs, A/B test readiness reviews, sample size and power calculations, launch checklists, rollout and rollback plans, GTM briefs, and go/no-go decision decks. Your job is to grade them against a rubric and explain your scores in writing. That means catching the errors a practitioner would catch — an experiment with no primary metric declared, a guardrail metric that can't actually be measured in the current instrumentation, a sample size estimate that ignores variance or novelty effects, a staged rollout that skips a kill-switch owner, or a launch deck whose narrative doesn't match the data on the slide.
Work typically arrives as batched tasks. For each one you score dimensions such as factual accuracy, methodological rigor, completeness against what a real launch review would demand, and presentation quality (slide hierarchy, chart honesty, readable tables). Written feedback matters as much as the score: reviewers who can say precisely which claim is unsupported and what the correct approach would be are the ones who get sustained task flow.
What the platform screens for
- Verifiable hands-on history. Expect follow-ups on specific launches or experiments you ran, your role in them, and what you'd change. Generic PM vocabulary doesn't survive probing.
- Experimentation literacy. Practical understanding of hypothesis framing, metric selection, MDE and power, peeking and stopping rules, segmentation traps, and when not to run a test at all.
- Evaluation judgment. Whether you can apply a rubric consistently, separate style disagreement from substantive error, and avoid both leniency drift and nitpicking.
- Writing and tooling. Clear structured English feedback, plus real fluency in Google Slides/PowerPoint, Sheets/Excel — you'll be judging deck and spreadsheet craft, not just prose.
Logistics
Fully remote, fully asynchronous, hourly. There are no set shifts; you claim tasks when available. Most contributors work in blocks of a few hours, and volume fluctuates with project demand — treat this as supplemental rather than guaranteed weekly load. Onboarding usually includes a calibration set with feedback before paid volume opens up. Pay bands are as observed on the platform and are set per project, not guaranteed.