Evaluation
From GLUE to agentic harnesses: benchmark lineages, the contamination crisis, and what independent measurement looks like now.
Evaluation has the shortest technique half-life in this guide: every benchmark generation was saturated, contaminated, or gamed within roughly two years, and the unit of measurement moved from the sentence to the task to the multi-hour session. The current regime’s defining feature is distrust.
The lineage
Figure 1. The evaluation lineage. The survivors are moving targets by design.
The three recurring failures
| Failure | Mechanism | Countermeasure |
|---|---|---|
| Saturation | Frontier passes the ceiling; benchmark stops discriminating | Harder sets, rotating content, expert-authored privates |
| Contamination | Test items leak into training corpora | Decontamination, post-cutoff items, parallel private forks (GSM1k) |
| Gaming | Optimization against the metric, not the ability | Style controls in arenas, held-out judges, methodology disclosure |
The agentic turn
- Task benchmarks (SWE-bench lineage) measure completion of real work, but scores are harness-dependent: the same weights move several points across scaffolds, and vendor tables mixing harnesses are not directly comparable (documented in our K3 audit).
- Verified subsets and containerized reproduction became the credibility bar after unverifiable early results.
- Long-horizon evals (multi-hour sessions, cost-per-completed-task) replaced single-response scoring at the frontier; cost joined accuracy as a first-class axis.
Independent measurement
Vendor-reported launch tables systematically exceed independent reproductions (endpoint load, harness choice, effort settings). The working stack in 2026: independent aggregators (Artificial Analysis-style indexes, arena Elo), plus first-party reproduction on your own workload before deployment. That last step is the only one that measures what you will actually ship; it is the method behind our model matrix and the audits on this site.
Ledger
| Verdict | Techniques |
|---|---|
| Dead | GLUE-era suites as frontier signal; static public test sets as sole evidence; single-number model comparisons |
| Current | Agentic task benchmarks with fixed harnesses, rotating/private sets, LLM-judge with bias controls, independent aggregators, cost-per-task reporting, safety and dangerous-capability evals, hallucination and abstention-aware scoring |
| Contested | LLM-judge reliability at frontier parity; contamination detection sensitivity; whether arena preference tracks task competence |
End of the guide. Start again at the Overview, or see the chronological companion: the LLM timeline.