What the work involves

This is not a build-something engagement — at least not yet. The first ask is your knowledge. Mercor wants engineers who have operated agents in production to walk through how those systems actually behaved: what broke, how you noticed, what you changed, and how you knew the change helped. The screen is conversational and untimed in spirit — no coding exercise, no take-home.

The topics of interest sit deliberately below the usual demo layer:

  • Internal monoagents wired into company data and permissions
  • Shared company memory, and what happens when it goes stale or wrong
  • Reusable skills and playbooks, and who maintains them
  • Tool and MCP surfaces agents call, including failure and retry behavior
  • Observability and cost attribution — how anyone actually sees what an agent did and what it spent
  • Long-lived execution: scheduled jobs, multi-hour work, cloud sandboxes
  • Adoption reality inside your org: who used the internal assistant, and who quietly stopped

What the platform screens for

The AI interview is looking for specificity, not vocabulary. Expect follow-ups that press on numbers, dates, and decisions: which eval you ran before and after a prompt change, what your regression rate looked like, why you chose a judge model over human review or the reverse. Candidates who describe agents in the abstract tend not to advance. Candidates who tell one messy story with real tradeoffs — including the ones that went badly — tend to.

Evaluation judgment matters as much as engineering. Mercor cares whether you can distinguish a real quality regression from noise, and whether you have opinions about what agent evaluation still cannot measure.

Logistics

Fully remote and asynchronous for the first stage; complete the AI screen whenever it suits you, typically 20–40 minutes. Strong screens are invited to a live 30-minute call with the Mercor team, scheduled to your availability. Compensation for that call has been observed in the $100–500 range, paid on completion, with the figure tied to depth of experience — not guaranteed, and set by the platform. Every submission is reviewed.