What the work actually involves

You will be handed survey instruments — some human-written, many AI-generated — and asked to judge whether they would produce decision-useful data. In practice that means going question by question for leading language, double-barreled items, unbalanced or overlapping response options, missing 'don't know' escapes, scale-point mismatches, and ordering or priming effects. It also means stepping back to the instrument level: does the screener actually admit the intended population, does the skip logic strand respondents, is the quota plan going to produce a sample anyone should generalise from, are there data-quality traps for straightlining and speeding.

A large share of tasks are comparative: two versions of a questionnaire, or two pieces of model reasoning about a design problem, and you pick the better one and say why. The 'why' is the deliverable. Reviewers who write "Q7 is leading" get flagged; reviewers who write "Q7 presupposes the respondent has already switched brands, which will inflate agreement among non-switchers and contaminate the segmentation cut" get kept. Many prompts are deliberately ambiguous — there is no single correct answer, and the platform is measuring whether you can reason to a defensible position and name the tradeoff you accepted.

What the screen looks for

Mercor's screening is AI-led: an application, a résumé parse, and a recorded conversational interview with adaptive follow-ups. It is checking that you have personally designed or fielded surveys rather than commissioned them, that you can name the specific error type behind a bad question instead of gesturing at 'bias', and that your judgment holds up when the interviewer pushes back on your answer. Expect to be asked about real projects with specifics — sample sizes, incidence rates, panel providers, what went wrong in field and what you changed.

Logistics

  • Remote, contractor, largely asynchronous; you claim tasks from a queue.
  • Volume fluctuates with project phase — some weeks are thin, some ask for real commitment. Most people are asked about a realistic weekly floor (often 10–20 hours).
  • Written English at professional standard is non-negotiable; the rationales are the product.
  • Pay is stated as observed at $120/hr and can vary by task type, tier, and project. Not guaranteed.