What the work involves

You sit between a research team and the realities of running a retail business. Day to day, that means authoring retail problems the model is likely to get wrong — assortment tradeoffs, markdown cadence, planogram and space productivity decisions, vendor funding and trade terms, allocation and replenishment logic, shrink and inventory accuracy, promo lift forecasting — and then writing the reference answer with reasoning that a director-level peer would sign off on. The other half is evaluation: reading model responses, scoring them against a rubric, and writing feedback that explains precisely where the reasoning broke, not just that it did.

Expect to also work upstream on the rubrics themselves. Retail judgment is rarely binary — two reasonable merchants can disagree on a buy — so a large part of the job is defining what "correct" means, where partial credit belongs, and how calibration should hold across multiple SMEs working the same task family.

What the platform screens for

  • Depth that survives follow-up. Mercor's screen probes specifics: which categories you owned, what your open-to-buy looked like, what the actual decision was and what it cost.
  • A real evaluation track record. Prior hands-on rubric-based scoring of LLM or AI outputs is stated as mandatory — describe it concretely, including the rubric structure and volume.
  • Career progression at a recognized retailer (Amazon, Walmart, Target, Costco, Home Depot, Nike, or equivalent), with 8+ years dedicated to retail.
  • Written clarity. Feedback that a research engineer with no retail background can act on.

Logistics

Remote and W-2 through Cincinnatus LLC as employer of record, with placement inside a leading AI lab's extended workforce. The listing asks for reliable engagement of at least 35 hours/week during weekdays — this is closer to a full-time schedule than typical async gig work, though the hours themselves are largely self-managed around syncs with research teams. Pay is observed in the $60–80/hr range and is not guaranteed; final rates depend on the screen, specialization, and lab placement.