What the work actually involves

This is not a software job in the usual sense — you are producing the material that teaches and tests models. Day-to-day work tends to fall into a few patterns depending on the project: authoring realistic backend tasks (a failing service, an N+1 query, a race condition in a job queue) with reference solutions and test cases; reviewing model-generated code and writing structured critiques of why an approach is wrong or merely mediocre; and ranking multiple model outputs against a rubric where the differences are subtle. Some projects lean toward multi-turn agentic evaluation, where you judge whether a model correctly navigated a real repository, read logs, and made a defensible change.

The hard part is rarely knowing the right answer. It is articulating why — in writing, precisely enough that another reviewer or a training pipeline can use it. Engineers who are fluent in tradeoff language (consistency vs. availability, normalization vs. read performance, retries vs. idempotency) do well here. Engineers who can only say "this looks fine" do not.

What the platform screens for

Mercor's process is application, credential verification, then an AI-led interview. The interview probes depth: expect to be asked about systems you personally built, then pushed on specifics — what the actual throughput was, why you chose that datastore, what broke in production and how you found it. Vague or resume-recited answers get filtered. The screen also assesses written communication and whether your stated availability is realistic, since project matching depends on it.

  • Verifiable professional backend experience — production systems, not coursework
  • Depth in at least one of: API/microservice design, database architecture and query optimization, distributed systems
  • Clear technical writing under time pressure

Logistics

Fully remote and largely asynchronous. Applying places you in a network, not a role — matching happens on a rolling basis when a lab requests backend expertise, and there can be a gap between approval and first invitation. Once matched, commitments typically run 15–30 hours weekly with flexible scheduling, though some projects impose review turnaround windows. The $70–150/hr band reflects observed rates across projects and varies with seniority, specialization, and the contracting lab; it is not a guarantee.