Ecosystem

Benchmark

Standardized evaluation datasets and metrics used to compare AI model performance across tasks like reasoning, coding, math, and language understanding.

Explained at five levels

Level 1

A test or quiz for AI to see how smart it is compared to other AIs.

Level 2

Standardized tests used to compare different AI models — like SATs for AI. They measure things like reasoning, coding, and knowledge.

Level 3

Standardized evaluation datasets and metrics used to compare AI model performance across tasks like reasoning, coding, math, and language understanding.

Level 4

Curated evaluation suites (MMLU, HumanEval, GSM8K, etc.) that measure model capabilities across defined tasks, enabling reproducible comparison but subject to contamination, overfitting, and construct validity concerns.

Level 5

Operationalized evaluation protocols measuring specific capability dimensions — subject to Goodhart's law, benchmark contamination via training data overlap, and the validity gap between benchmark performance and real-world task competence.

Definitions are educational summaries. Terminology can vary by source and context.

Sources