What the work actually involves
This is not a product engineering job. You are building the material that teaches and measures a model: writing realistic full-stack tasks (a broken auth flow, an N+1 query in a Django ORM layer, a React state bug that only appears on rehydration), producing reference solutions, and then judging model attempts against them. Much of the day-to-day is comparative grading — two model responses to the same prompt, and you explain in writing which is better and precisely why. Rubric work rewards engineers who can articulate why a solution is wrong, not just flag that it is. Some projects lean toward authoring multi-file repository tasks with test harnesses; others are pure evaluation with a heavy annotation quota.
What the platform screens for
Mercor's intake is resume verification plus an AI-conducted interview. The interview probes whether your claimed stack experience survives follow-up questions: expect to be asked to explain a tradeoff you made in production, then to defend it against a plausible counterargument. Screens tend to filter on shipped professional work across both tiers — a frontend framework (React, Angular, Vue), a backend framework (Node, Django, Spring, Rails), and real experience with both relational and NoSQL data models. Communication is weighted heavily because the deliverable is written reasoning; vague or hedging explanations read as shallow domain knowledge to an automated scorer.
Logistics
- Fully remote and largely asynchronous, with occasional calibration sessions or reviewer feedback loops.
- Typical commitments run 15–30 hours per week, contract-based, project by project.
- Applying joins a pool — matching is rolling, and there can be a gap between approval and a first project invitation.
- Observed pay for this network runs $70–150/hr; the actual rate is set per project and depends on task complexity and seniority, not guaranteed by the listing.
- Expect a second, project-specific interview before any engagement begins.