What the work involves
This is an evaluation seat, not a gig booking. Labs training multimodal and audio models need people who make a living playing to unpredictable crowds in unpredictable acoustics — because that's where model outputs break. Day to day you might:
- Record and annotate short performance clips (solo, ambient noise, crowd interruption, mid-song key changes) against a supplied rubric
- Rate model-generated accompaniment, transcription, or chord/rhythm detection for whether it would actually survive a real set
- Write prompts and adversarial cases: mistuned strings, dropped beats, improvised transitions, requests shouted from a crowd
- Compare two model responses on musical plausibility, timing, genre fidelity, and give a written justification a reviewer can audit
- Occasionally: live-session capture with defined mic and location conditions, or judgment calls on busking-adjacent topics (permits, crowd handling, set structure, tip dynamics)
What the screen looks for
micro1's screening is AI-led and follow-up heavy. It cares much less about a polished resume than about whether your specifics hold up. Expect to name instruments, keys, tunings, venues, pitch spots, and years played, then get pushed one or two layers deeper on each. Vague answers collapse quickly. The other half of the screen is evaluation judgment: can you separate "I don't like this" from "this is wrong," apply someone else's rubric faithfully, and flag ambiguity instead of guessing.
Logistics
Remote and mostly asynchronous. Batches arrive with deadlines rather than shifts; typical committed volume is 5–20 hours per week, and some projects want a minimum weekly floor. Audio-capture tasks require a quiet space plus a usable recording setup (interface or decent USB mic — phone-only is usually rejected). Rates in the $100–300/hr range have been observed on this listing; actual offers vary by project, are set per-engagement, and are not guaranteed.