What the work involves
This is evaluation work, not client design work. Instead of producing a shipped feature, you are judging model output: a generated onboarding flow, a set of component variants, a written design critique, a heuristic analysis of a screenshot, or a rationale explaining why one layout beats another. Your job is to say whether the output would survive a real design review — and then say precisely why, in language a model can learn from. Typical tasks include scoring two candidate responses against a rubric, rewriting a weak model answer into a reference-quality one, flagging where a model invents plausible-sounding design reasoning (fake accessibility claims, misused Nielsen heuristics, invented platform guidelines), and writing new prompts that expose those failures.
Because the listing text micro1 published for this role is essentially empty, treat the specifics as project-dependent. In practice these engagements cluster around UI critique, design-system reasoning, accessibility judgment, and multimodal tasks where the model is shown a screen and asked to assess or improve it.
What the platform screens for
- Verifiable shipped work. Named products, your actual scope on them, and what changed after you designed it. Portfolio links matter more than titles.
- Depth under follow-up. micro1's AI interviewer asks a design question, then pushes: why that spacing, why that pattern, what did you rule out. Shallow answers collapse in two turns.
- Written precision. Evaluation output is prose. Vague critique ("feels cluttered") is unusable; specific, falsifiable critique is the product.
- Rubric discipline. Willingness to score against someone else's criteria rather than your own taste, and to flag when the rubric itself is wrong.
Logistics
Fully remote and largely asynchronous, with work delivered through a task queue. Most specialists treat it as part-time — commonly 10–20 hours per week — with volume that fluctuates as projects open and close. Expect an onboarding calibration round where your early scores are compared against reviewers before full volume unlocks. You need your own machine, a reliable connection, and working access to the design tools you claim (Figma above all).