The work

You receive artifacts an AI model produced in response to an L&D brief: a 30-60-90 onboarding plan, a facilitator guide for a manager training workshop, a skills matrix, a slide deck for a compliance module, a spreadsheet tracking completion and assessment scores. Your job is to score each one against a rubric and write feedback that a model trainer can act on. That means separating surface polish from instructional substance — a deck can look clean and still sequence objectives backward, mislabel Kirkpatrick levels, invent a citation for a learning-science claim, or build knowledge checks that test recall when the stated objective was application.

Most tasks run 30–90 minutes. You will be asked to justify scores, not just assign them, and to flag specifics: broken formulas, inconsistent slide masters, objectives written without measurable verbs, activities with no time allocation, accessibility problems in a deck meant for a global rollout.

What the screen looks for

  • Verifiable depth. Expect follow-ups on programs you have actually built — audience, duration, delivery mode, how you measured whether it worked.
  • Rubric discipline. Whether you can hold a consistent standard across artifacts and resist rewarding fluent writing that is instructionally hollow.
  • Real tool fluency. Slides and PowerPoint especially; also spreadsheet mechanics, since many artifacts are trackers or budget models.
  • Feedback quality. Concrete, structured, and specific enough that someone who is not an L&D professional understands the defect.

Logistics

Remote and asynchronous. Hourly, with work drawn from a queue rather than assigned on a schedule — volume fluctuates with project demand, and steady weeks are not guaranteed. Most contributors work 10–20 hours a week. The $80–120/hr band reflects rates observed on this platform for expert evaluation work in this category; your offer depends on assessed depth and the specific project.