What the work involves
You receive tasks — prompts, reference solutions, test harnesses, or model outputs — built around multi-service AWS serverless applications, and you decide whether they hold up. That means reading CDK, CloudFormation, or SAM templates closely enough to spot a resource policy that grants more than it should, an EventBridge rule whose pattern will never match, a Step Functions retry block that swallows a permanent failure, or a DynamoDB Streams consumer that quietly assumes exactly-once delivery. Some batches ask you to grade candidate solutions against a rubric; others ask you to audit the task itself — is the problem statement unambiguous, is the reference solution actually correct, does the test pass for the right reason?
The written feedback carries most of the weight. A verdict without a traceable reason is not usable by the lab, so each judgment needs to name the service behaviour or IAM semantic that makes the artifact wrong, and where relevant, what a correct version would do instead.
What the platform screens for
- Depth under follow-up. Mercor's AI screen asks a question, then pushes on the answer. Naming Step Functions is not the same as explaining the difference between a Standard and Express workflow's execution history under failure.
- Real-AWS versus emulator judgment. Familiarity with LocalStack matters less than knowing where its fidelity breaks — IAM enforcement, eventual consistency, service quotas, stream ordering.
- Calibration. Can you distinguish a solution that is stylistically unlike yours from one that is genuinely wrong, and hold a consistent bar across a batch?
- Verifiable history. Specific systems you built, specific stacks you wrote, specific incidents you debugged.
Logistics
Remote and asynchronous. Work arrives in batches with turnaround windows rather than fixed shifts; most reviewers commit 10–20 hours a week and are paid hourly against reviewed time. Observed rates for this listing sit in the $70–90/hr range, which varies with track, batch difficulty, and calibration performance — not guaranteed. Expect a paid or unpaid calibration set before full volume opens.