What the work actually involves

You spend most of your time in two modes. In the evaluation mode, you read AI-generated code — sometimes a function, sometimes a multi-file change with a diff and a stated design rationale — and judge it for correctness, quality, and design soundness. That means running it where possible, finding the input that breaks it, and writing feedback specific enough that a training pipeline can use it: not "the error handling is weak" but "swallows the IOError on line 34, so a truncated read returns success with partial data."

In the authoring mode, you build problems designed to be hard for a strong model — tasks with a real bug behind a plausible-looking surface, or system design prompts where the obvious answer fails under a constraint you introduce. Each comes with a reference solution and a test suite that actually distinguishes a correct answer from a lucky one. You will also review model reasoning traces on architecture, debugging, and system design questions, where the answer may be right for the wrong reasons — and saying so precisely is the whole job.

What the platform screens for

micro1 runs an AI-led interview before human review. It probes verifiable specifics of your engineering history — which company, which system, what you personally owned — and then follows up on depth. Expect to be asked to debug or critique something live, in words. The screen is looking for engineers who can articulate why a design choice is wrong, at a level of concreteness that reads as usable feedback rather than taste. Writing quality matters here in a narrow sense: precision, not polish.

Logistics

  • Contractor engagement, fully remote, asynchronous task queues
  • Hours are flexible; many contributors work part-time alongside a full-time role
  • Preference stated for candidates based in an English-speaking country
  • Preferred background: at least a year as an engineer at a well-known technology company, with some of that experience in the last seven years
  • Pay band is as observed on the platform for this role and varies by task complexity and calibration performance; it is not a guarantee