What the work actually involves
You will spend most of your time doing two things. First, building task-specific grading criteria for real brand design briefs — identity systems and guidelines, logos, campaign and marketing visuals, packaging concepts, presentation and layout work, digital and social assets, and brand messaging and positioning. A usable rubric here is not "good typography, 1–5"; it specifies what a passing hierarchy looks like for this brief, what a scalable mark must survive, where a positioning statement fails against the stated audience.
Second, you score completed work samples — some AI-generated, some human — against those criteria and write the justification for every score. The written argument is the deliverable. Reviewers are checking whether a peer reading your justification could arrive at the same number without you in the room. Expect senior-reviewer feedback, calibration against other designers, and fast iteration cycles.
What the platform screens for
- Agency provenance, specifically. The brief names Pentagram, Wolff Olins, Landor, Collins, IDEO as the reference class. In-house and enterprise design teams are explicitly not the target profile. Expect your resume and portfolio to be read for studio names, roles, and client work.
- Craft depth under follow-up. Screens push past vocabulary. You should be able to explain why a particular type pairing fails at small sizes, or how you'd diagnose a logo that reads well at poster scale and dies as a favicon.
- Critique that sounds like a creative director in review — precise, prioritized, and separable from personal taste.
- Consistency. Two similar samples should get similar scores from you, with the reasoning tracing back to the rubric rather than to instinct.
Logistics
Fully remote and asynchronous, contract engagement through Mercor. Observed rates for this posting run $80–150/hr, set at onboarding based on background and calibration performance — as observed, not guaranteed. Task volume on evaluation projects fluctuates; most contributors treat it as part-time (roughly 10–20 hours a week) alongside client work. Copywriting, messaging, and prior AI-evaluation experience are listed as pluses, not requirements.