The Center for AI Standards and Innovation (CAISI) is a US federal AI standards body within the National Institute of Standards and Technology (NIST), Department of Commerce. It is the successor to the AI Safety Institute (AISI) created under Biden Executive Order 14110; it was renamed and rescoped under the Trump Administration in 2025 to emphasize "standards and innovation" over "safety." CAISI produces guidance, workshops, and publications on AI evaluation, monitoring, and standards, and convenes the NIST AI Consortium with industry and federal-agency stakeholders.
Activities
CAISI's work spans workshops and consortia, technical publications, and standards crosswalks. In 2025 it held three workshops on post-deployment AI system monitoring with external stakeholders, federal agencies, and the larger NIST AI Consortium, which draws subject-matter experts from more than 10 federal agencies and roughly 200 external experts in academia and industry. It produces NIST AI publications on monitoring, evaluation, agent standards, and risk management.
CAISI has published crosswalk documents mapping the NIST AI RMF to ISO/IEC 42001, the OECD AI Principles, and other frameworks. Standards Are the New Legislation (Clearwater, March 2026) describes these crosswalks as a "Rosetta Stone" for global governance interoperability that most observers overlooked.
Publications
- NIST AI Risk Management Framework — predates CAISI but is maintained by NIST.
- NIST AI 600-1 Generative AI Profile.
- NIST AI Agent Standards (2026).
- NIST AI 800-4 Challenges to the Monitoring of Deployed AI Systems (March 2026), described as CAISI's most extensive substantive report on monitoring.
Evaluation work
CAISI has issued four reports finding Chinese models significantly less secure than US counterparts (Source: washingtonpost.com). Its published evaluation of DeepSeek R1-0528 found that the model complied with 94% of jailbreak prompts, against 8% for US reference frontier models. Anthropic cites this benchmark in its "2028: Two Scenarios" policy paper as evidence for tightening export controls and disrupting distillation pathways. The result has been used as an anchor for arguments about a safety gap between US and Chinese frontier models.
AI Agent Standards Initiative
CAISI launched the AI Agent Standards Initiative on February 17, 2026, organized around three pillars: industry-led standards (voluntary guidelines feeding US leadership in international standards bodies); pre-deployment evaluation infrastructure (voluntary benchmarking access ahead of public release); and critical-infrastructure deployment guidance (coordination with CISA on sector-by-sector deployment). See NIST AI Agent Standards Initiative (Feb 2026).
Pre-deployment evaluation and the institutional-home question
On May 5, 2026, CAISI signed pre-deployment evaluation agreements with Microsoft, Google, and xAI, establishing a multi-lab voluntary benchmarking channel for frontier models ahead of release. On May 11, the Commerce Department deleted the May 5 announcement from its website, reportedly at the urging of the White House Office of National Cyber Director (ONCD). Reporting framed the deletion as a sign that the ONCD prefers a different institutional home for frontier-model pre-release evaluation, with some accounts pointing to an ODNI-led center as the favored alternative and Commerce opposing that arrangement; whether the deletion reflected one-off political housekeeping or a broader withdrawal of CAISI from the pre-deployment-vetting role was not clear from the reporting (Source: reuters.com).
Bipartisan signals on May 12-13, 2026 — a 32-member House letter to NCD Cairncross and an open letter from ICBA, BSA, and TechNet — both rejected mandatory pre-release review and favored ONCD-as-coordinator, positions that align with CAISI's voluntary-track role even as the May 5 announcement was removed.
OpenAI published its frontier safety blueprint on June 2, 2026, and summarized it in a Global Affairs post the following day. The blueprint calls for CAISI to become the federal government's primary frontier-AI safety authority, with permanent authorization and funding, flexible hiring, and classified-compute access. It frames CAISI as an independent evaluator rather than a deployment-approving regulator: it would certify third-party assessors and track progress toward recursive self-improvement, but would not hold a pre-clearance or licensing role. OpenAI's proposal aligns with the voluntary-track posture reinforced by the May 2026 signals and with the "visibility, not approval" architecture of the June 2 executive order, and it endorses CAISI over an ODNI-led alternative as the institutional home for frontier evaluation (Source: openaiglobalaffairs.substack.com). See AI Pre-Release Vetting.
On June 9, 2026, Trump administration officials — including National Cyber Director Sean Cairncross — directed CAISI to halt publication of its model assessments while the June 2 executive order is implemented, a pause reported to reflect a push by Cairncross and Treasury Secretary Scott Bessent to give national-security considerations a larger role in how AI models are evaluated. The move threw the unit's public-facing future into doubt despite praise for its work from AI developers, and bears on whether independent evaluators publish assessments of newly released frontier models such as Claude Fable 5 (Source: wsj.com).
CAISI researchers tested and endorsed the new safety classifier Anthropic deployed for its July 1, 2026 redeployment of Fable 5 after the export-control lift, an operational evaluation role during a period when the unit's public assessments remained paused (Source: anthropic.com; thehackernews.com).
CAISI's testing was also the stated basis for the Commerce Department's July 7, 2026 clearance of OpenAI's GPT-5.6 for broad public release, ending the staggered rollout imposed in June; OpenAI announced public launch of the Sol, Terra, and Luna tiers for July 9 (Source: axios.com).
OpenAI restated and sharpened its institutional preference on August 3, 2026, saying it had pushed the administration and bipartisan members of Congress for CAISI to hold a central role in determining which frontier systems undergo review, what criteria apply, and how quickly reviews happen — an allocation of the gatekeeping question to CAISI that goes further than the June 2 blueprint's certify-and-track framing, while keeping the evaluate-don't-license boundary (Keeping America out in front on AI (OpenAI Global Affairs, August 2026)).
Leadership
Chris Fall, appointed CAISI director in April 2026, resigned on July 20, 2026, three months into the role; NIST Director Arvind Raman became acting director while Commerce seeks a permanent replacement (Source: axios.com; theinformation.com).
Funding
Axios reported on July 10, 2026 that CAISI operates on a $15 million budget against an estimated $84 million annual need, a gap cited in accounts of the administration's reliance on ad hoc, negotiation-based model vetting rather than standardized severity frameworks (Source: axios.com). See AI Pre-Release Vetting.
Fathom CEO Andrew Freedman argued in a July 1, 2026 Transformer op-ed that the Supreme Court's June 29 Slaughter ruling — which held the president may remove independent-agency heads at will — strengthens the case for accredited independent verification organizations reporting to CAISI rather than a new independent agency, noting Connecticut and Virginia have already passed independent-verification bills (Source: transformernews.ai).
Standards as governance baseline
Per Andrew Clearwater's analysis (Standards Are the New Legislation (Clearwater, March 2026)), NIST's frameworks are increasingly the de facto national governance baseline for US AI. Clearwater identifies several channels: Texas TRAIGA cites NIST frameworks for safe-harbor compliance on an incentive lane; Illinois SB 3312, Utah HB 286, and Washington HB 2157 cite them on a mandate lane; California TFAIA / SB 53 uses them on a transparency lane for required disclosure; and courts use them to define standard-of-care in negligence and strict-liability cases even where no statute mandates it. Clearwater compares the pattern to the trajectory of the NIST Cybersecurity Framework and notes bipartisan support for leveraging NIST to develop technical AI standards.
Proposed statutory design
OpenAI's June 2026 blueprint proposes authorizing CAISI "as a permanent institution with clear statutory authorities and sufficient funding to conduct frontier model evaluations, develop safety standards, certify third-party assessors, and coordinate with national" partners — while confining its role to evaluation and recommendation: "CAISI's role should be to conduct evaluations and recommend mitigations — not to approve or block deployments." It further proposes a statutory evaluation deadline, after which "if CAISI fails to complete an evaluation within the defined time period due to bandwidth, hardware, personnel, or other constraints, developers should be permitted to deploy without penalty." This is the sharpest structural divergence from Anthropic's framework, which contemplates an Agency able to fine, prohibit further deployment, and in extreme cases restrict access to deployed models.
Alongside authorization and appropriation, the blueprint sets out four further foundational measures. The CAISI Director "should report directly to the US Secretary of Commerce or another senior Cabinet-level official," with White House coordination of staffing and operational support across departments. Hiring authorities "similar to those used by CHIPS for America" would permit rapid technical recruitment at competitive compensation, alongside temporary tours of duty for outside researchers — a proposal that speaks to the $15 million budget against an estimated $84 million need reported above. National security and scientific agencies would be directed to "immediately make available personnel and data related to cyber, CBRN, and other national security domains." And CAISI would obtain "access to classified computing environments capable of conducting frontier model evaluations," through dedicated infrastructure, interagency partnerships, or agreements with agencies or commercial providers.
The blueprint would also make CAISI the certifying body for an assessor market rather than the sole assessor. It asks CAISI to develop standards for independent technical assessments, "establish a certification process for qualified third-party assessors," and require frontier developers above specified capability thresholds to undergo periodic assessments by certified organizations, with findings reported to CAISI. To build that market, CAISI would be authorized "to provide grants, cooperative agreements, and other forms of support to emerging assessment organizations, academic centers, and technical institutions" — a construction that bears on the accredited-verifier proposals discussed above (see Independent Verification Organizations (IVOs)). Stated first priorities for these assessments are recursive self-improvement progress, "highly capable internal deployments, frontier model security, internal monitoring practices, and the effectiveness of associated safeguards" (see Rogue Internal Deployment).
Two August 2026 proposals from outside industry assign CAISI different roles. The IFP report of August 6, 2026 makes CAISI the implementing body for seven of its 23 recommendations, and treats resourcing as the binding constraint: Congress should fund it at a minimum of $84 million per year — the figure IFP had recommended previously, alongside a staff of 184, and against the America First Policy Institute's recommendation of $50–100 million per year — with the authors describing this as a minimum bar rather than a target. They argue autonomy and business model matter more than size, proposing a direct line of communication to a senior White House or cabinet-level official, authority to forward-deploy staff into frontier AI companies to increase visibility into major research projects and internal deployments, authority to organize third-party evaluators and data producers, and authority to contract directly with frontier developers to support priority research and security work, including compute subsidies for alignment teams. Further recommendations would have CAISI develop guidelines for managing the risks of rapid AI capability improvement and co-lead an AI Verification Consortium with industry.
ARI's August 10, 2026 blueprint gives CAISI a narrower, advisory role and places regulatory authority elsewhere. CAISI supplies technical assistance toward a workable capabilities-based definition of a frontier model, to replace the compute threshold as that proxy degrades, and contributes technical analysis to each standards-setting cycle; the assurance and enforcement functions sit with a regulator the blueprint does not fix, offering the Department of Commerce as one option and a distribution across several agencies as another.
Measuring recursive self-improvement is the blueprint's most specific technical ask of the agency. It states that policymakers "currently have limited visibility into RSI progress, whether safeguards are keeping pace, or what indicators should inform future policy decisions," directs CAISI to work with developers, academics, national security agencies and international partners "to rapidly develop methodologies, benchmarks, and indicators for measuring RSI," and adds: "we encourage CAISI to treat RSI as an urgent priority." See Recursive Self-Improvement (RSI).
- related: How Should the US Prepare for Increasingly Automated AI R&D? (IFP, August 2026) — names CAISI in seven of 23 recommendations and proposes an $84M/184-staff floor
- related: Responsible Innovation at the Frontier (ARI, August 2026) — assigns CAISI technical assistance on a capabilities-based frontier definition
Relationships
- part-of: US Department of Commerce, NIST
- predecessor: AI Safety Institute (AISI; pre-2025 rename)
- maintains: NIST AI Risk Management Framework 1.0
- publishes: NIST AI 800-4: Challenges to the Monitoring of Deployed AI Systems and other NIST AI series
- related: UK AI Safety Institute (AI Security Institute), EU AI Office (international counterparts)
- related: Post-Deployment AI System Monitoring, Standards as Litigation Evidence