An AI benchmark is a snapshot of particular tasks at a particular moment. Left unchanged, it can encourage optimisation for old examples while real use moves elsewhere. Maintaining the evaluation set requires rules that refresh it without losing comparability.
Keep a stable core
Maintain a set of representative cases that changes infrequently and provides a historical reference. Include ordinary tasks, critical failures and difficult exceptions. The core enables comparisons between versions without constantly moving the target.
Add a rolling layer
Bring in new material from real, appropriately protected cases, tickets and human corrections. Include more than failures: retain the natural distribution of use. The set then tracks reality without becoming a collection of extremes.
Define refresh triggers
A new use category, policy change, model, language or recurring failure warrants reconsideration. Record who decides on additions and which benchmark version accompanies each release.
Prevent leakage into the prompt
If exact test answers are continually used for optimisation, the score stops demonstrating generalisation. Restrict access to part of the set, create variations and check new cases. Separate development from final evaluation.
Improve the judge as well
Scoring instructions should describe correct, partially correct and dangerous answers. Check agreement between human reviewers on a sample and reconsider ambiguous criteria. An unstable rubric can create more noise than the model itself.
Publish scores with a version
Every result should include the set version, date, model, settings and known limitations. Do not compare figures from different benchmarks as if they were the same measurement. Transparency in evaluation is part of quality.
The living AI benchmark is an original evaluation framework developed by DigitalNow.