What the work involves
Most days you are producing or reviewing infrastructure-flavored evaluation data. That can mean authoring a realistic task — a Helm chart that drifts from its values file, a Terraform plan that silently destroys a stateful resource, a GitHub Actions pipeline that passes locally and fails in CI — then writing the reference solution and the rubric a grader will apply. Other days you sit on the review side: comparing two model attempts at a Kubernetes debugging trace, ranking them, and writing the justification that explains which one an on-call engineer could actually act on.
The judgment being measured is specific. Models are fluent about YAML and cloud APIs and frequently produce output that reads correctly while being operationally wrong — deprecated API versions, IAM policies that are broader than stated, health checks that mask a real failure, retries that amplify an outage. Your value is catching those and articulating the failure mode in a sentence a non-specialist reviewer can verify.
What the platform screens for
micro1 runs an AI-led interview before any project match. Expect it to probe depth rather than breadth: you will be asked to talk through concrete systems you have built or broken, and follow-ups will push on the details a resume can't fake — what the rollout strategy was, why you picked that state backend, what the postmortem concluded. Communication clarity matters as much as technical range, because the deliverable is usually written reasoning. Candidates who describe generic "DevOps best practices" without specifics tend not to advance.
Logistics
- Fully remote and predominantly async; work is claimed from a queue rather than scheduled
- Part-time is normal — many contributors commit 10–20 hours per week alongside a full-time role
- Observed rates for this category run $60–130/hr, varying with specialization and review responsibility; not guaranteed
- Engagements are contract-based, often starting with a paid calibration batch before steady volume