What the work involves
This is an expert-sourcing pipeline for search engineering judgment, not a generic coding gig. Tasks in this category typically center on evaluating and stress-testing agentic retrieval behavior: reading a query and an agent's tool-call trace, judging whether the retrieval step was the right call, deciding whether a result set actually answers the user's intent, writing rubrics for relevance, and constructing hard cases that expose failure modes like query rewriting drift, over-fetching, stale indexes, or reranker overconfidence. Observed pay is $80–150 per task, which varies with task length and complexity and is not guaranteed.
The screening interview itself is 25 minutes and conversational. Mercor is listening for specifics: what system you owned, what metric you moved, how you knew it moved, and what you gave up to move it. Vague ownership claims fall apart under follow-up, so the interviewer will press on details — offline versus online metric disagreements, how you built eval sets, what your ground truth was, why NDCG or recall@k was or wasn't the right lens.
What the platform screens for
- Direct ownership of relevance, retrieval, or ranking on a system with real users
- A working theory of measurement: how you quantified quality when labels were scarce or noisy
- Hands-on exposure to agentic search — embeddings plus tool use, multi-hop retrieval, query planning, RAG pipelines in production
- Data infrastructure fluency: indexing, freshness, throughput, cost tradeoffs at scale
- Ability to explain a messy tradeoff clearly and without polish
Logistics
Fully remote and asynchronous. The interview is self-scheduled and takes 25 minutes with no coding component and no take-home. Standout candidates are invited to a paid 30-minute live conversation with the team, observed at $200 paid on completion of the call. Ongoing task work is contract-based with self-selected volume; most contributors treat it as part-time alongside a primary role.