What the work actually involves

You spend most of your day producing and judging infrastructure problems that are hard enough to expose where a model's reasoning breaks down. That means authoring scenarios — a cluster where pods are stuck in `CrashLoopBackOff` for a reason that isn't in the logs, a Terraform plan that applies cleanly but creates a drift trap, a Lambda–API Gateway–DynamoDB path with a throttling failure mode — then writing the reference solution with enough diagnostic detail that a reviewer can follow every inference. The other half is evaluation: reading model output or another expert's submission, deciding whether the diagnosis was correct for the right reasons, and writing feedback that names the specific error rather than gesturing at it.

You will also be asked to build the rubrics. Turning "good IaC design" into criteria a second engineer can apply and reach the same score on is the part most people underestimate, and it is where consistency across the dataset comes from.

What the screen looks for

Mercor's screen is AI-led and follows up. It pushes hard on whether your Kubernetes experience is operational or declarative — authoring manifests against a managed control plane is explicitly not what this listing wants. Expect to be asked about a specific cluster incident: what you saw first, what you ruled out, what the root cause turned out to be. The same applies to Terraform and CDK (state management, module boundaries, blast radius) and to AWS integration work you personally shipped. Vague seniority claims collapse quickly under a second question; concrete incidents with named components hold up.

Logistics

  • Remote, W-2 employment with Cincinnatus LLC as employer of record, placed within a client AI lab team
  • 40 hours/week, weekday-aligned — this is not a nights-and-weekends side engagement
  • Heavy written output; async collaboration with other SMEs on rubric calibration
  • Observed rate $75–110/hr, typically set by depth of production experience