What the work involves

You receive artifacts an AI model produced in response to a people ops or recruiting brief: a leveling framework in a spreadsheet, a hiring manager intake doc, a DEI reporting deck, a structured interview guide, a PIP template, a total-rewards summary for a board meeting. Your job is to decide whether it would survive contact with a real HR leader. That means checking factual accuracy (does the model understand FLSA exempt tests, at-will language, I-9 timelines, FMLA thresholds?), methodological rigor (is the comp benchmarking logic coherent? are the funnel metrics computed correctly?), and presentation quality (does the slide deck actually read cleanly, or is it dense text on a broken layout?).

You then write structured feedback: what is wrong, why it matters, and what a competent practitioner would have produced instead. Tasks typically arrive in batches with a rubric attached; scoring dimensions vary by project but usually separate correctness from style, and ask you to justify each score in writing.

What the platform screens for

  • Specific, verifiable experience. Mercor's screen goes deep on what you actually owned — req volume, headcount supported, systems you administered, whether you wrote policy or executed it. Vague "HR generalist" framing tends not to clear.
  • Depth under follow-up. Expect the interviewer to push a second and third question on the same topic: you say you built a leveling framework, it asks how you handled compression and how you priced against survey data.
  • Evaluation judgment. Can you separate "this is wrong" from "this is not how I'd do it"? Evaluators who penalize stylistic preference as error are the most common failure mode.
  • Deck and spreadsheet fluency. Slides and Sheets/Excel proficiency is a hard requirement, not a nice-to-have — a meaningful share of artifacts are presentation-format.

Logistics

Fully remote and asynchronous. Work is hourly and drawn from a task queue, so volume fluctuates with project demand rather than following a fixed schedule; most evaluators treat it as part-time supplementary work. There is usually a calibration period where your scores are compared against reviewer consensus before you're given steady volume. Pay bands are as observed on the platform for this category and are not guaranteed for any individual engagement.