What the work actually involves
Most shifts are a queue of audio clips, usually 5 to 60 seconds, with a task template attached. You might tag instrumentation and production characteristics on a library of stems, mark the onset of a key change, transcribe lyrics with disfluencies and ad-libs preserved, or write a caption dense enough that a text-to-audio model could plausibly reconstruct the clip from it. A second stream of work is comparative: two model-generated renditions of the same prompt, and you decide which is better and say why in terms an engineer can act on — "B has a smeared transient on the snare and the vocal sits behind the mix" rather than "B sounds worse."
You will also spend real time on edge cases: clips where the genre label is genuinely contested, clips with two instruments that occupy the same register, clips where the reference prompt asked for something the model didn't deliver at all. Flagging these clearly, with a reason, is often worth more to the client than pushing volume.
What the platform screens for
micro1's screen is AI-led and conversational, and it follows up. Expect to be asked to distinguish things that sound similar — a Rhodes from a Wurlitzer, a plate reverb from a hall, a compression artifact from a bad room — and to explain how you'd hear the difference rather than simply assert it. Expect at least one guideline-conflict scenario, because rubric adherence under ambiguity is the trait that actually separates hires here. Vocabulary matters: you should be comfortable in both musician's terms and audio-engineering terms, and able to switch when a rubric demands one or the other.
Logistics
- Remote, contractor, invoiced hourly against tracked task time.
- Asynchronous; work is pulled from a queue, with occasional calibration sessions scheduled with notice.
- Most contributors are asked for a minimum weekly commitment (commonly 10–20 hours) with some consistency week to week.
- You supply your own monitoring. Closed-back consumer earbuds are generally not sufficient for artifact-detection tasks; studio headphones and a quiet room are the practical baseline.