What the work actually looks like
This is not a single posting — it's an open application to a talent network. Once verified, you're eligible for contracts that come in on a rolling basis from AI labs. The underlying work varies by project, but for DevOps and platform engineering it usually falls into a few shapes:
- Task authoring: writing realistic infrastructure problems — a broken Helm chart, a Terraform state conflict, a CI pipeline that passes locally and fails in the runner — with a verifiable correct resolution.
- Model output evaluation: reading what a model produced for a pipeline config, IaC module, or incident triage prompt, and judging whether it would actually work in a real cluster or just looks plausible.
- Comparative ranking and rubric feedback: choosing between two model responses and writing the reasoning trail that explains why, in terms a non-expert annotator could not have produced.
The hard part is rarely knowing the answer. It's articulating why an answer is wrong — misapplied IAM scoping, a race condition in a rollout strategy, a Dockerfile that builds but bloats or leaks secrets — with enough precision that the feedback is usable as training signal.
What the platform screens for
Mercor's intake is an AI-conducted interview after resume upload and location confirmation. It probes depth rather than breadth: expect follow-ups that push past your first answer on cloud primitives, Kubernetes internals, CI/CD design tradeoffs, and observability. Vague or credential-adjacent answers tend to stall. Reviewers are looking for someone who has personally owned production infrastructure — debugged a bad deploy at 2am, migrated a cluster, argued about blue-green versus canary with real constraints — not someone who has read about it. Written clarity is weighted heavily, because the deliverable is usually prose.
Logistics
Fully remote and largely asynchronous. Typical engagements run 15–30 hours per week, and project duration varies from a few weeks to several months. You choose which invitations to accept. Observed rates for this category sit between $70 and $150 per hour, depending on the client, the specialization required, and your assessed depth — this is a range seen on the platform, not a guarantee. Note the structural reality: acceptance into the network does not mean immediate work, and match timing depends on lab demand.