What the work actually involves

This is not a product engineering job. You are building the material that teaches and measures a model: writing realistic full-stack tasks (a broken auth flow, an N+1 query in a Django ORM layer, a React state bug that only appears on rehydration), producing reference solutions, and then judging model attempts against them. Much of the day-to-day is comparative grading — two model responses to the same prompt, and you explain in writing which is better and precisely why. Rubric work rewards engineers who can articulate why a solution is wrong, not just flag that it is. Some projects lean toward authoring multi-file repository tasks with test harnesses; others are pure evaluation with a heavy annotation quota.

What the platform screens for

Mercor's intake is resume verification plus an AI-conducted interview. The interview probes whether your claimed stack experience survives follow-up questions: expect to be asked to explain a tradeoff you made in production, then to defend it against a plausible counterargument. Screens tend to filter on shipped professional work across both tiers — a frontend framework (React, Angular, Vue), a backend framework (Node, Django, Spring, Rails), and real experience with both relational and NoSQL data models. Communication is weighted heavily because the deliverable is written reasoning; vague or hedging explanations read as shallow domain knowledge to an automated scorer.

Logistics

  • Fully remote and largely asynchronous, with occasional calibration sessions or reviewer feedback loops.
  • Typical commitments run 15–30 hours per week, contract-based, project by project.
  • Applying joins a pool — matching is rolling, and there can be a gap between approval and a first project invitation.
  • Observed pay for this network runs $70–150/hr; the actual rate is set per project and depends on task complexity and seniority, not guaranteed by the listing.
  • Expect a second, project-specific interview before any engagement begins.