What the work actually involves

Mercor is building high-fidelity replicas of a marketing org's tool surface — CRM, product analytics, project tracker, CMS, ads, web analytics, support queue — populated with realistic documents, dashboards, personas and deliberately planted problems. Your job is to supply the product marketing half of that: task briefs drawn from launches and positioning work you have personally owned, the tool state a realistic version of that task requires, and the reference answer that separates a strong practitioner from a plausible-but-wrong one.

A typical unit of work starts as a brief — say, a launch two weeks out where the pricing page copy contradicts the sales deck, adoption data in Amplitude disagrees with the CRM opportunity notes, or a competitor's response invalidates half the positioning. You specify the artifacts that make that scenario real, state the planted issue explicitly, and write the criteria a grader applies. Then you review AI agent attempts against your own standard and explain, in writing, why a given answer misses.

What the platform screens for

  • Hands-on ownership, not oversight. The screen probes for launches you ran, positioning you wrote, enablement you shipped — with named artifacts and outcomes. Team-management framing reads as a gap here.
  • Tool fluency under follow-up. Expect specific questions about what you pulled from Salesforce, Amplitude, Asana or Linear, and how you reconciled them when they disagreed.
  • Evaluation judgment. Can you articulate the difference between an answer that sounds like a PMM and one that is actually correct? That distinction is the product.
  • Written reasoning. Much of the deliverable is prose explaining why a decision is right. Terse or hedged writing is a real disqualifier.

Applicants complete a short multiple-choice knowledge screener specific to product marketing before human review. Prior experience building AI training environments, RL environments or simulated case studies is strongly preferred, as is Rubric Academy or Rubric Bootcamp certification — neither is stated as mandatory.

Logistics

Fully remote and asynchronous, contract-based, with work claimed in units rather than assigned by shift. Observed pay for this band runs $60–100/hr, set by depth of experience and screener performance — stated as observed, not guaranteed. Contributors typically commit 10–20 hours a week; the environment-design work is bursty, and review passes on agent attempts arrive in batches with turnaround expectations attached.