What the work involves

You build self-contained HR "worlds" — a realistic organizational setting, a task an AI system is asked to perform inside it, a reference output showing how a senior practitioner would actually handle it, and a rubric that separates real judgment from plausible-sounding recall. Scenarios span workforce planning, recruiting and talent acquisition, employee and labor relations investigations, performance management, compensation and benefits design, and organizational restructuring. Many require you to reconstruct the tooling context that HR leaders actually work in: HRIS platforms such as Workday or SAP SuccessFactors, applicant tracking systems, and compensation benchmarking sources like Radford or Mercer surveys.

The reference artifacts are the ones you have produced in-role — an investigation report with findings and remediation, a compensation analysis with band construction and defensible pay-equity logic, a RIF or reorg proposal with adverse-impact review, a policy rewrite triggered by a legal change. Rubrics are the harder half of the job: you have to articulate why one answer is defensible and another is exam-quality boilerplate, and write that distinction so a reviewer who is not you can apply it consistently.

What the platform screens for

  • Ownership, not adjacency. Mercor's screen probes matters you personally ran — the investigation you led, the comp cycle you designed, the reorg you defended — with follow-ups that reward specifics and expose secondhand knowledge quickly.
  • Legal fluency in practice. Not statute recitation, but how Title VII disparate-impact analysis, FLSA exemption tests, ADA interactive-process obligations, or EU consultation requirements actually constrain a decision under time pressure.
  • Track clarity. Be explicit about US versus International; jurisdictional bluffing is the fastest disqualifier.
  • Writing under a rubric mindset. Prior rubric, training-content, or policy authorship is a meaningful advantage.

Logistics and pay

Fully remote and asynchronous, contractor engagement, with contributors observed in the $70–80/hr band for this listing — rates are as reported, not guaranteed, and vary with track, credential, and calibration performance. Most people work 10–20 hours per week around a full-time role, with periodic synchronous calibration sessions. Expect an onboarding period where early tasks are reviewed closely before volume opens up; SHRM-SCP, SPHR, or an international equivalent is strongly preferred.