Ecosystem
Benchmark
Standardized evaluation datasets and metrics used to compare AI model performance across tasks like reasoning, coding, math, and language understanding.
Explained at five levels
Level 1
A test or quiz for AI to see how smart it is compared to other AIs.
Level 2
Standardized tests used to compare different AI models — like SATs for AI. They measure things like reasoning, coding, and knowledge.
Level 3
Standardized evaluation datasets and metrics used to compare AI model performance across tasks like reasoning, coding, math, and language understanding.
Level 4
Curated evaluation suites (MMLU, HumanEval, GSM8K, etc.) that measure model capabilities across defined tasks, enabling reproducible comparison but subject to contamination, overfitting, and construct validity concerns.
Level 5
Operationalized evaluation protocols measuring specific capability dimensions — subject to Goodhart's law, benchmark contamination via training data overlap, and the validity gap between benchmark performance and real-world task competence.
Definitions are educational summaries. Terminology can vary by source and context.
Sources
- NIST Trustworthy and Responsible AI Resource Center: Glossary — National Institute of Standards and Technology. Accessed 2026-07-20.