You don't need to attend ML conferences to benefit from their program lists. For evaluators, they're a leading indicator of what labs will pay experts to do next.
The recurring pattern
Papers announce that frontier models have saturated an existing benchmark; within months, harder replacements appear — built from questions written and verified by credentialed specialists, because crowd-sourced questions no longer challenge frontier models.
Why this matters to evaluators
- Benchmark construction is expert work. When a lab or academic group builds a new domain benchmark, they contract specialists to write and verify items — short, well-paid engagements that surface on expert platforms.
- Grading is moving from answers to reasoning. Recent evaluation research focuses on judging chains of reasoning and agent behavior, not final answers — exactly the skill platforms now screen for.
- LLM-as-judge has documented limits. A steady stream of papers shows model graders diverging from expert graders precisely in specialized domains. Human expert review remains the ground truth there.
How to position yourself
Practice two skills: writing hard questions in your specialty (the kind a strong generalist gets wrong), and grading reasoning step-by-step against a rubric. Both map directly onto where lab budgets are heading.