What the work actually involves
You build and judge tasks that look like real developer work rather than textbook exercises. That means seeding a repository with a plausible commit history, wiring a GitHub Actions workflow that fails for a specific reason, configuring an OAuth flow against Slack or Google Workspace, then watching an AI agent attempt the task and grading what it produced. Grading is hands-on: you run the code, read the CI logs, and check whether the diff actually does what the agent claims. A second strand of the work is exploratory — probing undocumented behaviour in tools and APIs and writing up what you find so the same environment can be reproduced by someone else.
Expect calibration sessions where several contributors grade the same output and reconcile disagreements against a rubric. Your written justifications matter as much as your verdicts; a score with no reproducible evidence behind it is of little use downstream.
What the screen looks for
- Concrete Git/GitHub fluency — branching strategy, review habits, reading diffs, diagnosing a failing pipeline from logs alone
- Ability to author CI workflows, not just consume them, and to make a build reproducible on a clean runner
- Real integration experience: REST APIs, token and OAuth flows, webhooks, and the failure modes each has
- Discipline under a rubric — whether you can hold a consistent line across many similar cases and flag ambiguity instead of guessing
- Clear technical writing in English, since environments and findings are handed to other people
No AI or ML background is required. Familiarity with agent tool-calling or MCP is a genuine plus but not a gate.
Logistics
Fully remote contractor engagement, largely asynchronous, with periodic live calibration calls. Volume typically fluctuates by project phase, so candidates who can commit predictable weekly hours tend to get steadier assignment flow. Pay in the $50–70/hr band has been observed for this role; rates are set per project and are not guaranteed.