What the work actually involves

This is engineering work first and evaluation work second. You are handed shifting requirements and early access to model APIs, and you ship real applications, services, and internal tools on top of them — backend and API design, a React-or-equivalent front end, data modelling, and cloud deployment. Around that you build the scaffolding the research team needs: tool interfaces, evaluation harnesses, telemetry. The valuable output is twofold: working code, and precise written accounts of how the model failed while you were building with it — where tool calls degraded, where long-context reasoning broke, where the API contract did not survive contact with production.

The cohort is staffed deliberately across language ecosystems. Deep production ownership in at least one of Python, Java, Rust, C#, or C++ is the bar, plus real delivery in a second language from a different ecosystem. Expect to work directly with the AI lab's engineering manager, onboard onto unfamiliar codebases fast, and produce working increments without much up-front specification.

What the screening measures

Mercor's screen is AI-led and drills on verifiable specifics. Expect follow-ups that push past framework names into architecture decisions, tradeoffs you rejected, incident history, and what you personally owned versus what your team owned. Also assessed:

  • Breadth claims held up under probing — a second language you can genuinely defend, not one you touched once
  • Ramp-up evidence: a concrete story of landing meaningful code in an unfamiliar repo within days
  • Written communication, since failure-mode reports are a primary deliverable
  • Career progression at recognized engineering organizations, 6+ years of shipped production work

Logistics

Remote, weekday-aligned, and genuinely full-time — 40 hours per week with reliable overlap for collaboration with the engineering manager. This is W-2 employment with Cincinnatus LLC acting as employer of record, including payroll, benefits, and compliance, with placement inside the client lab's extended workforce. Pay in the $90–110/hr band reflects observed rates for this cohort and is not guaranteed; final rate depends on language depth, seniority, and the lab's staffing needs.