You don't need to attend ML conferences to benefit from their program lists. For evaluators, they're a leading indicator of what labs will pay experts to do next.

The recurring pattern

Papers announce that frontier models have saturated an existing benchmark; within months, harder replacements appear — built from questions written and verified by credentialed specialists, because crowd-sourced questions no longer challenge frontier models.

Why this matters to evaluators

  • Benchmark construction is expert work. When a lab or academic group builds a new domain benchmark, they contract specialists to write and verify items — short, well-paid engagements that surface on expert platforms.
  • Grading is moving from answers to reasoning. Recent evaluation research focuses on judging chains of reasoning and agent behavior, not final answers — exactly the skill platforms now screen for.
  • LLM-as-judge has documented limits. A steady stream of papers shows model graders diverging from expert graders precisely in specialized domains. Human expert review remains the ground truth there.

How to position yourself

Practice two skills: writing hard questions in your specialty (the kind a strong generalist gets wrong), and grading reasoning step-by-step against a rubric. Both map directly onto where lab budgets are heading.