What the work looks like
You are handed a short video clip — often only a few seconds — plus a written directive such as "pick up the red mug and place it on the shelf without knocking over the bottle." Your job is to decide whether the recorded behaviour actually satisfies that directive, and to say precisely where it doesn't. That means marking action boundaries frame-accurately, writing descriptions that a model can learn from, and flagging the moment a grasp slipped, an object was nudged, or a sequence was performed out of order. A second common task type is pairwise comparison: two clips, same instruction, which one better fulfils it and why. A third is repair work — an AI system has already produced annotations, and you tighten the boundaries, fix wrong object references, and delete hallucinated events.
What the platform screens for
micro1's screen is conversational and AI-led. It probes whether you can hold a definition of "task complete" steady across dozens of ambiguous clips, whether you notice small differences in movement and object contact rather than judging on overall impression, and whether you can write a one-line justification that another annotator would reach the same conclusion from. Expect follow-ups that push on a judgement you just made — the screen is testing whether your reasoning survives pressure or collapses into "it just looked wrong." Prior AI experience is explicitly not required; relevant domain grounding includes robotics, manufacturing or warehouse operations, kinesiology and movement analysis, physical or occupational therapy, sports video review, film or broadcast editing, ergonomics, and prior annotation or QA work.
Logistics
- Fully remote contractor engagement, project-based rather than salaried.
- Largely asynchronous: batches of clips with rubric documents, worked on your own schedule, with periodic calibration rounds or guideline updates.
- Observed pay band for this listing is $15–80/hr, stated as observed and not guaranteed; placement within it tracks task complexity, calibration performance, and the depth of domain background you bring.
- A reliable computer and connection matter more than most people expect — frame-level scrubbing on long batches is unpleasant on a weak setup.