What the work involves
This is engineering work, not annotation. You'll be writing production-quality code — services, internal tools, front ends, evaluation harnesses — on top of models that haven't shipped yet. A typical week mixes building working increments against loosely specified requirements with the surrounding scaffolding that makes frontier models usable: tool interfaces, telemetry, retry and error handling, and test rigs.
The second half of the job is written. When a pre-release model breaks in your integration — bad tool calls, silent format drift, reasoning that collapses on long context — you're expected to isolate it, reproduce it, and write it up clearly enough that a research team can act on it. Engineers who treat the write-up as an afterthought tend not to last in these cohorts.
What the platform screens for
- Verifiable production history. 3+ years shipping software at a recognized organization, with specifics: what you owned, what scale, what broke.
- Language breadth. Production depth in at least one of Python, Java, Rust, C#, or C++, plus real working proficiency in a second. The cohort is deliberately staffed across ecosystems.
- Genuine full-stack range. Backend APIs, a modern front-end framework, relational or document data stores, and cloud deployment you actually operated.
- Fast ramp on unfamiliar code. Expect questions about the last time you shipped inside a codebase you didn't write.
- Written clarity. Screens typically include a written component; terse or vague answers read as a signal.
Logistics
Fully remote from India, full-time at 40 hours/week during weekdays, with overlap expectations set by the lab's engineering manager. Engagements are contract-based and scoped by project, so duration varies with the lab's roadmap. Pay in the $29–35/hr range has been observed for this listing; actual offers depend on cohort, language track, and interview outcome, and are not guaranteed.