NIST AI 200-2, "The TEVV-Athlon Framework for Evaluating AI Systems," is a draft framework published by NIST as an initial public draft announced on August 7, 2026. It describes a method for developing customised assessments of AI systems starting from an organisation's own test, evaluation, verification and validation (TEVV) objectives, rather than supplying a fixed benchmark suite (NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026)).
| Field | Value |
|---|---|
| Designation | NIST AI 200-2 (initial public draft) |
| Issuing body | National Institute of Standards and Technology, U.S. Department of Commerce |
| Series | NIST Trustworthy and Responsible AI |
| Authors | P. Jonathon Phillips, Theodore Jensen, Patrick Hall, Razvan Amironesei, Yee-Yin Choong, Craig Greenberg, Kristen K. Greene |
| Announced | August 7, 2026 (cover date August 2026) |
| DOI | 10.6028/NIST.AI.200-2.ipd |
| Comment period | 60 days; opened August 7, 2026, closes October 6, 2026 |
| Comments to | TEVV-Athlon@nist.gov, subject line "NIST AI 200-2" |
| Status | Draft |
Scope
The draft takes its definition of TEVV from the final report of the National Security Commission on Artificial Intelligence: "A framework for assessing, incorporating methods and metrics to determine that a technology or system satisfactorily meets its design specifications and requirements, and that it is sufficient for its intended use." NIST states the framework is intended to cover statistical machine learning models, large language models, multi-modal models and agentic systems — a scope that spans model classes treated separately in most existing evaluation guidance (NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026)).
The draft describes its design constraint as a "Goldilocks" problem: a universal framework is too unwieldy to apply, a specialised one too narrow to generalise. Its stated aim is to balance comprehensiveness and flexibility, and it carries the epigraph "Homage to George Box: No evaluation framework is perfect, some are useful."
Structure
The framework sets out four stages for constructing an assessment from an organisation's TEVV objectives. The product of that process is termed a TEVV-Athlon: an assessment in which systems are tested through a set of Events and Tools that generate data on Blocks, the measurement concepts of interest. The draft distinguishes the two terms explicitly — the TEVV-Athlon Framework is the methodology, a TEVV-Athlon is the assessment it produces.
| Stage | Purpose |
|---|---|
| Articulate & Organize | State the goal; select the system attribute or trustworthiness characteristic to be evaluated; account for lifecycle stage. |
| Define & Construct | Establish the Blocks — "the key concepts or metrics which the evaluation aims to assess." |
| Apply & Measure | Determine the Events ("the activities which produce evidence for each Block") and assemble the Toolbox ("set of methods and instruments used to elicit, collect, and analyze information from the Events"). |
| Synthesize & Interrogate | Analyse the collected data by working back through the stages in reverse, and present results. |
Stage 1 supplies seven questions for articulating goals, presented as building on the Heilmeier Catechism and the Feynman method, beginning with a statement of objectives "using absolutely no jargon" and covering audience, timeline and cost, resources, prior techniques built on, the basis for success, and the challenges (NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026)).
A worked example applies the framework to what the draft calls the query-violation problem — a chatbot expected to answer a request while withholding a prohibited category of information — across three scenarios (TV shows, meal planning, travel planning), adapted from the NIST ARIA 0.1 pilot evaluation. It measures the trustworthiness characteristics Valid and Safe through the Blocks Helpfulness and Violation Frequency, with User Testing and Red Teaming as Events.
Section 4 supplies general guidance on assembling expertise, checking resources, composing a toolbox from model testing and from red-teaming, user testing and field testing, applying scientific measurement practices, validating that a measurement captures what it intends, and guarding against Goodhart's Law. The draft states plainly that benchmarks "often measure proxies contained in datasets rather than real-world outcomes," that open benchmarks are exposed to leakage, contamination and saturation, and that "strong benchmark performance or metric scores do not always indicate broader system quality or suitability for deployment." Five appendices tabulate selected benchmarks, red-teaming methods, user and field testing methods, measurement-science considerations, and experimental-design considerations, each keyed to trustworthiness characteristics and lifecycle stages.
Comment process
NIST invites input on any aspect of the draft and enumerates six areas: the definitions and uses of the TEVV terms; the framework's flexibility, scope and applicability; types of TEVV process or activity the framework may not adequately address; its usefulness for developing new TEVV processes and for novel or emerging systems; aspects warranting clarification, revision, removal or expansion; and additional concepts, provisions, examples or supporting materials. Comments are subject to release under the Freedom of Information Act, and NIST states its staff may use software tools including AI to summarise or analyse comments while submitted data will not be used to train models (NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026)).
Reception
Ike Harris, executive director of the Frontier Security Institute, described the proposal as "the first step in standardizing the way the federal government evaluates AI systems both for itself and for its contractors" (Source: reuters.com). That characterisation reaches beyond the draft's stated scope, which addresses organisations generally rather than federal use specifically; the draft does not itself claim procurement application.
The draft appeared the same day President Trump said Congress wants to regulate the AI industry "out of business," a juxtaposition drawn directly by the reporting (Source: reuters.com).
Relation to other frameworks
NIST AI 200-2 addresses how to construct an evaluation, where the AI Risk Management Framework addresses how to organise risk management around one, and NIST AI 300-1 addresses how to document datasets and models for public release. The draft states its own place in that division directly: a TEVV-Athlon implements the AI RMF's Measure function, the Govern and Map functions supply its inputs at Stage 1, and its Stage 4 results feed the Manage function. It cross-references the Center for AI Standards and Innovation's Practices for Automated Benchmark Evaluations of Language Models (NIST AI 800-2 ipd) for LLM benchmarking guidance (NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026)). Its subject overlaps with the evaluation-design questions raised by the third-party assessment schemes proposed by frontier developers (Making AI Audits and Assessments Work (OpenAI Global Affairs, August 2026)) and by the classified federal benchmarking process described under AI Pre-Release Vetting, though the draft does not reference either.
Relationships
- depends-on: NIST AI 200-2 ipd — The TEVV-Athlon Framework for Evaluating AI Systems (initial public draft, August 2026) — the primary text.
- related: NIST AI Risk Management Framework 1.0, NIST AI 300-1 — Guidance and Templates for Public-Facing AI Documentation, NIST AI Agent Standards Initiative (2026) — the adjacent NIST AI publications.
- related: AI Benchmarks and Evaluation — the measurement problem the framework addresses.
- related: AI Pre-Release Vetting, Government AI Procurement.
- depends-on: National Institute of Standards and Technology (NIST) — the issuing body.