Deployment

Evaluation

From GLUE to agentic harnesses: benchmark lineages, the contamination crisis, and what independent measurement looks like now.

Evaluation has the shortest technique half-life in this guide: every benchmark generation was saturated, contaminated, or gamed within roughly two years, and the unit of measurement moved from the sentence to the task to the multi-hour session. The current regime’s defining feature is distrust.

The lineage

Figure 1. The evaluation lineage. The survivors are moving targets by design.

The three recurring failures

FailureMechanismCountermeasure
SaturationFrontier passes the ceiling; benchmark stops discriminatingHarder sets, rotating content, expert-authored privates
ContaminationTest items leak into training corporaDecontamination, post-cutoff items, parallel private forks (GSM1k)
GamingOptimization against the metric, not the abilityStyle controls in arenas, held-out judges, methodology disclosure

The agentic turn

  • Task benchmarks (SWE-bench lineage) measure completion of real work, but scores are harness-dependent: the same weights move several points across scaffolds, and vendor tables mixing harnesses are not directly comparable (documented in our K3 audit).
  • Verified subsets and containerized reproduction became the credibility bar after unverifiable early results.
  • Long-horizon evals (multi-hour sessions, cost-per-completed-task) replaced single-response scoring at the frontier; cost joined accuracy as a first-class axis.

Independent measurement

Vendor-reported launch tables systematically exceed independent reproductions (endpoint load, harness choice, effort settings). The working stack in 2026: independent aggregators (Artificial Analysis-style indexes, arena Elo), plus first-party reproduction on your own workload before deployment. That last step is the only one that measures what you will actually ship; it is the method behind our model matrix and the audits on this site.

Ledger

VerdictTechniques
DeadGLUE-era suites as frontier signal; static public test sets as sole evidence; single-number model comparisons
CurrentAgentic task benchmarks with fixed harnesses, rotating/private sets, LLM-judge with bias controls, independent aggregators, cost-per-task reporting, safety and dangerous-capability evals, hallucination and abstention-aware scoring
ContestedLLM-judge reliability at frontier parity; contamination detection sensitivity; whether arena preference tracks task competence

End of the guide. Start again at the Overview, or see the chronological companion: the LLM timeline.