What the work actually involves

You will spend most of your time reading code you did not write and deciding whether it is correct, and then proving your decision. Typical tasks include authoring realistic engineering problems with hidden test suites, writing the reference solution yourself, comparing two model responses and justifying a preference in writing, and repairing prompts where the model was right for the wrong reason. Some projects lean toward agentic evaluation — running a model through a multi-step task in a real repo and annotating where its plan degraded. Others are narrower: unit-test authoring, docstring fidelity, or diff review in a single language.

What the platform screens for

micro1 runs an AI-led interview before any human sees your profile. It probes stated experience for specificity — which language, which version, what broke, how you found it — and it follows up when an answer stays general. Expect live technical questioning rather than a take-home: naming the concurrency primitive you used and why you rejected the alternative counts for more than describing a project's business impact. The second thing being measured is evaluation judgment: whether you can articulate a defect precisely, distinguish style disagreement from a real bug, and hold a consistent standard across similar submissions.

Logistics

  • Fully remote, contractor engagement, invoiced hourly against tracked or delivered work.
  • Mostly asynchronous, though some projects require overlap with US business hours for calibration calls.
  • Commitments commonly range from 10 hours a week to full-time; project volume fluctuates and gaps between batches are normal.
  • You need your own development machine, a stable connection for the interview, and the ability to work in English.

Pay bands are as observed on the platform, not guaranteed. Rate placement typically tracks language rarity, systems depth, and how cleanly you performed in calibration.