What the work actually looks like
This is an operations seat on a data project, not a strategy seat. On a typical day you're pulling annotator throughput and quality numbers into Google Sheets, spotting the contributors whose accuracy is drifting, sampling their recent submissions to figure out why, and writing the message that corrects it. You'll field contributor questions about ambiguous task instructions, escalate genuine spec gaps to the research or client side, and keep a running picture of whether the project will hit its volume target on schedule.
The subject matter is LLM browsing — tasks where a model navigates the live web to answer questions or complete multi-step goals. You don't need to have built one, but you do need to be able to read a task rubric closely enough to tell a real annotator error from a rubric ambiguity, because that distinction drives most of your decisions.
What the screen looks for
- Concrete ownership stories. Expect follow-ups on scale: how many contributors, what volume, what the quality metric actually was and how it moved.
- Spreadsheet fluency under questioning. Not "proficient in Excel" — how you'd structure a tracker, what formulas or pivots you'd reach for, how you'd catch a data-entry problem before it reaches a report.
- Evaluation judgment. Scenarios where two annotators disagree, or where a top performer's numbers look too good, and what you'd do next.
- Written communication. Much of the role is asynchronous messaging to contributors who may be in other time zones; the screen will probe tone and clarity when delivering corrective feedback.
Logistics
Remote and largely asynchronous, though PM roles on Mercor tend to carry more real-time overlap expectations than annotation work — some projects want a few hours of predictable daily availability for contributor escalations. Engagements are commonly structured as ongoing part-time or near-full-time contract work tied to a specific project's lifespan. Be honest in the screen about your weekly hours and time zone; underdelivering on stated availability is the most common way these engagements end early.