AI Policy Wiki
Dashboard

Gap Scan — 2026-07-29

Daily gap hunt — what the wiki is missing and how each gap was triaged.

Scanned

Recent window: 2 dev-log files in the last 48h (2026-07-28 08:05 and 22:05); 60 wiki pages with last_updated inside 48h. Rotation slice: 13 — sources/ Q–Z (122 pages) plus the Raw Sources/ orphan check (744 raw files; 21 flagged as name-not-mentioned, all confirmed false positives — supporting sources folded as (Source: URL), whose filenames were never intended to appear in wiki text).

Broken-link scan across 1,685 content pages: 345 distinct broken targets. Alias resolution run on the top 120 by in-degree; 3 resolved to existing pages and were fixed, and a fourth (in-degree 1, below the sweep cut) was caught by the post-edit verification pass, the rest are either genuine missing-page candidates below threshold, escaped-pipe artifacts of the known bin/lint-scan.py capture bug, or intentional Raw Sources/ citation defects in lint's lane.

Gaps actioned (7 of 24 found)

New pages created (live)

Pages expanded (live)

  • Safety Cases for Frontier AI (score ~5) — the slice's worst reliance/depth ratio. In-degree 22 against 325 words summarizing a 5,393-word raw source, with sources_count absent and last_updated 2026-06-06. The prior page carried the four-component definition and four bullet claims; it named none of the paper's substance. Deepened to roughly 3,400 words directly from Raw Sources/Safety Cases for Frontier AI.md — the primary text, not the web. Added: the arXiv date (28 October 2024) and Buhl as corresponding author; the paper's four explicit scope restrictions and its definition of catastrophic risk; the MoD (2007) definition and the three-stage produce/review/decide process; the safety-case history in energy, petrochemicals and transportation and the rules-based approach it displaced after the 1960s–80s accidents, together with the recorded criticism (false assurance, little empirical evidence for efficacy, practitioners consider them effective, UK recognised best practice); the full four-component table with each component's stated requirement; the illustrative scope block (one trillion parameters, 10²⁶ FLOP pre-training on Common Crawl, RLHF, free-account API access, monitoring assistant) and the capability-buffer assumption (one year of scaffolding/prompting gains plus >10% of pre-training compute); the ≥10⁻⁷/year of ≥1,000 fatalities risk threshold, the three-region ALARP structure, and the autonomous-vehicle comparative objective; the three argument families from the European Rail Agency's risk-acceptance principles and the four-way Clymer et al. (2024) taxonomy (inability, control, trustworthiness, deference); the five capability red lines in the worked sketch and the authors' disclaimer that the sketch is illustrative; the evidence taxonomy and the mechanistic-understanding turn; the safety-framework relation including the 16-companies-committed / four-published figure and the three roles cases play relative to frameworks; the six-dimension rules-versus-cases comparison in both directions; the six preferability conditions with the paper's likely/unclear/possible assessment of each; the lifecycle treatment; all five implementation challenges; and both policy-recommendation sets in full. Frontmatter repaired (source_type date corrected from bare 2024 to 2024-10-28, sources_count 1, tags extended). Confidence stays medium — single-source summary of one paper, below the three-source bar. Every prior fact, citation and typed relationship preserved; original backed up to _meta/_revision-backups/2026-07-29/.

Corrections (live)

  • EU AI Act (Regulation 2024/1689) (score ~6) — a live-thread correction to a page carrying a two-way date disagreement that has now become three-way. The page recorded the Digital Omnibus on AI as entering into force "in the week of 10 July 2026," with a second paragraph noting later reporting placing it in August 2026 (signed 8 July, Official Journal publication by 24 July, in force twenty days after). Practitioner commentary published 2026-07-28 supplies a third account: the AI Omnibus Regulation entered into force 27 July 2026. Folded into the existing "accounts differ" paragraph as a third reading rather than replacing either — none of the three is independently confirmable from the primary Official Journal record the wiki holds, and the substantive deadlines (2 December 2027 standalone; 2 August 2028 product-embedded) are undisputed across all three. Also added the two open EDPB drafts bearing on generative AI under the GDPR rather than the AI Act: draft Guidelines 02/2026 on anonymization, which replace 2014 Article 29 Working Party guidance and shift the test from absolute identifiability to likelihood of identification by a given entity, and draft Guidelines 03/2026 on web scraping in the context of generative AI, adopted 7 July 2026 and open for consultation until 30 October 2026. Frontmatter: last_updated 2026-07-29, sources_count recounted 25 → 16 from distinct inline URLs (the prior figure was not derivable from the page). [2 new (Source: URL) cites]

Frontmatter repaired (live)

  • sources_count added to the 8 slice-13 sources/ pages that lacked it: stanford-hai-ai-index-2026, science-of-scheming, safety-cases-frontier-ai, taxonomy-systemic-risks-gpai, shape-of-the-thing, the-bitter-lesson, standardization-trends-ai-safety, scaling-era-ch1 — all resolved to 1 (the anchoring raw source, with no distinct inline (Source: URL) citations). last_updated deliberately not bumped on the seven that were not otherwise edited. Slice 13 now has zero pages missing sources_count and zero missing source_type.

Queued — foundational sources

  • "Pacing the Frontier" — statement from employees of frontier AI companies (2026-07-28) (score ~8: live thread +3, dangling foundational +3, wiki core area +2) — gap type 1. Zero occurrences of the phrase anywhere in Wiki/. A substantive coalition letter — the category CLAUDE.md treats as nearly always foundational — signed at release by 1,178 frontier-lab employees including the chief scientist or chief science officer of OpenAI, Anthropic, Meta AI and Thinking Machines, Anthropic's CEO and two co-founders, and Google DeepMind's chief strategy officer. Verified: pacingthefrontier.com (HTTP 200; og:title, title and og:url all match the request URL; og:description "A statement from over 1000 employees of frontier AI companies"; body dateline "JULY 2026"; three distinctive passages present verbatim including the request sentence; independently corroborated by the 2026-07-28 22:05 dev-log, whose separately-scraped signatory list is a subset of the site's own rendered roster). Signatory count is a moving figure — 1,178 at release, 1,224 when observed today — and the raw file plus the ingest task both carry the standing instruction to cite the count with its observation date, per the discipline the 07-27 scan established for the Open Weights letter. Saved: Raw Sources/Pacing the Frontier - Statement from Employees of Frontier AI Companies.md. Queued: INGEST-pacing-the-frontier-statement-2026-07-29.md. Verification trail: queue/gap-scan/proposed-sources/pacing-the-frontier-statement-2026.md.
  • Hugging Face, "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident" (2026-07-27) (score ~8) — gap type 1. Two source pages already summarize the July 2026 intrusion disclosures; neither is this document, and the slug agent-intrusion-technical-timeline appears nowhere in Wiki/. This is the forensic reconstruction: ~17,600 recovered agent actions in ~6,280 clusters between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC, with named exploitation vectors, per-phase and per-day action tables, the credential findings, and the defensive-asymmetry reflection. Verified: huggingface.co/blog (the company's own publication surface; og:url matches exactly; og:title matches; twitter:site @huggingface; four named Hugging Face authors with on-domain profiles; an "Update on GitHub" link to the post's source file in github.com/huggingface/blog, a second integrity signal on the same organization's infrastructure; four distinctive passages verbatim; both of its cross-references resolve to documents the wiki independently holds; independently corroborated by the 2026-07-28 22:05 dev-log's own scrape). Raw file deliberately not saved — at roughly 10,000 words plus code blocks and tables, transcribing it risks silent drift in exactly the timestamps and action counts that make it worth ingesting, so the INGEST task carries the verified canonical URL for a fresh fetch, per the gap-identifier rule for long documents. The full text was read end-to-end during verification and the ingest task records the substance to capture. Queued: INGEST-hugging-face-agent-intrusion-technical-timeline-2026-07-29.md. Verification trail: queue/gap-scan/proposed-sources/hugging-face-agent-intrusion-technical-timeline-2026.md.

Authenticity-verification failures

None this run. Both pulls verified on their canonical hosts with independent corroboration.

Finding for the ingest lane

Probable mischaracterisation on two already-ingested pages. The Hugging Face technical timeline's TL;DR calls ExploitGym "an OpenAI cyber-capability evaluation harness," but its own body names SunBlaze-UCB/exploitgym and links the Berkeley repository, a commenter states plainly that it is a Berkeley RDI benchmark from Dawn Song's team rather than an OpenAI harness, and Dawn Song's signatory comment on the Pacing the Frontier statement independently claims CyberGym and ExploitGym as her own evaluation work. Three independent signals against the looser formulation. Security Incident Disclosure — July 2026 (Hugging Face) and OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation (OpenAI, July 2026) should be checked for it. Not corrected here — the correction depends on the queued ingest, and rewriting two ingested source pages from a third document's comment thread is outside gap-identifier's lane.

Deferred backlog (over the daily cap — re-surfaces next run)

  • Slice-13 residue, in reliance order. Techno-Federalism: How Regulatory Fragmentation Shapes the U.S.-China AI Race (in-deg 89, 1,124w against a 33,339-word raw file — a 30:1 ratio and the slice's largest absolute mismatch, ~5; a source-robustness-check job rather than a gap-scan side pass); Frontier AI Safety Commitments (Seoul, 2024) (in-deg 57, but 916w against a 910-word raw — proportionate, no gap); TAKE IT DOWN Act — Source Summary (in-deg 43, 627w/693w — proportionate); Stanford HAI AI Index Report 2026 (in-deg 37, 1,094w against a 136,054-word raw file — the worst ratio in the wiki by raw size, but the raw is a full annual index report and a proportionate summary is arguably correct; frontmatter repaired today, ~4); Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (in-deg 30, 535w/476w — proportionate); We Need a Science of Scheming (in-deg 28, 676w/2,929w, ~3); California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) (in-deg 26, ~3); Texas Responsible AI Governance Act (TRAIGA / HB 149) — Source Summary (in-deg 24, ~3); A Taxonomy of Systemic Risks from General-Purpose AI (in-deg 17, 825w against 11,376w, ~4). Note that unlike slices 11 and 12, most of the high-in-degree slice-13 pages are proportionate to their raw sources — the reliance/depth problem is concentrated in four pages, not the whole slice.
  • The high-on-sources_count: 1 question, now measured on a third consecutive slice. 78 of 122 slice-13 sources/ pages — 64% — carry confidence: high on sources_count: 1, against 65% on slice 12 (07-28) and 57% on slice 11 (07-27). Three slices at 57–65% settle it: this is an inherited ingest convention, not a set of individual errors, and the fix is a one-time policy decision for the curator or lint — either the decay table's "3+ sources" bar means something different for a sources/ page faithfully summarizing one primary document, in which case CLAUDE.md should say so explicitly, or roughly 60% of every sources/ slice is overrated. Still not actioned; a wrong call propagates to 558 pages.
  • bin/lint-scan.py target-capture bug — fourth consecutive run flagging it. Diagnosed 07-26, confirmed 07-27 and 07-28, confirmed again today: of the top 60 broken-link candidates, 16 are escaped-pipe artifacts (companies/anthropic\, companies/openai\, gpt-53-codex\, claude-mythos-5\, eu-ai-act\, california-sb-243\ and similar) and 4 more are the literal token wikilinks from prose about the link syntax. Proposed fix unchanged (\[\[([^\]|#]+?)(?:\\?\|[^\]]*)?\]\], or strip a trailing backslash from the captured target). Still unapplied; still lint's lane. Until it is fixed, roughly a third of the reported broken-link count is noise, which is the single largest source of false candidates in this scan.
  • Two citation-format defects, third consecutive flag. [[Raw Sources/raine-vs-openai-et-al-complaint.md]] on Raine v. OpenAI, Inc. and Raine v. OpenAI — Wrongful Death Complaint (2025); [[Raw Sources/Garcia-v-Character-Technologies-Inc Complaint.md]] on Garcia v. Character Technologies — Wrongful Death Complaint (2024) and Garcia v. Character Technologies, Inc.. CLAUDE.md's citation format never wikilinks a raw file; these should be (Source: Raw Sources/…). Four instances across four pages. Also [[_meta/briefings/weekly-2026-W19]] from three mainspace pages, which points from content into _meta/. Assigned to lint on 07-28 and left there; mechanical enough that a curator could clear all seven in one pass.
  • Recent-window items behind tonight's fold, not gaps yet. The Modal Labs confirmation that a customer asset was compromised in the same campaign, with Akshat Bubna's statement that the customer "published an unauthenticated endpoint" and that Modal "was not compromised in any way," plus OpenAI's same-day narrowing to four accounts across four services and its statement that no models planned for upcoming release were involved — this materially narrows what OpenAI currently records. Sam Altman's 2026-07-28 statement that the incident forced OpenAI to pause model training. Zuckerberg's WSJ op-ed ("invention, not automation, will be the greatest contribution of superintelligence"; closing models "will only disadvantage the US and its allies") and his FT position that the US should not ban Chinese AI — both paywalled below the lede, and the op-ed is a foundational source-queue candidate at ~5 once retrievable. The WSJ report of a Silicon Valley backlash against Anthropic over product-launch competition, data-retention changes and its open-weights position, with named quotes from Figma's Dylan Field and Notion's Sarah Sachs (a Anthropic fold). The Meta–BlackRock El Paso venture (Snapshot rows on Meta AI: 80/20 ownership split, ~$1B one-time distribution to Meta, $12.5B debt financing, ~$14B total development cost, 1 GW, capacity from 2028). ChatGPT approaching 1 billion weekly active users — not reported as crossed, and the publisher's hedge must survive the fold. Senate Commerce AI markup amendments on state-law preemption, lede-only.
  • entities/guidelight-ai-standards — new-page candidate at in-degree 0, surfaced by today's Pacing the Frontier pull as one of the two nonprofits giving the statement organizational support. Below threshold now; will rise once the statement is ingested. Check entities/encode-ai exists at that point.
  • Genuine missing-page candidates confirmed by alias resolution but below the in-degree threshold (2 each), carried forward. concepts/compute-thresholds is now built and comes off this list. Remaining: concepts/economic-possibilities-for-artificial-intelligence, concepts/algorithmic-decisionmaking, concepts/ai-and-language-models, concepts/talent-flow-china-us, concepts/data-broker-regulation, concepts/a-vision-of-democratic-ai, concepts/safety-training-methodologies, concepts/ai-coarse-grainings, concepts/cross-national-ai-policy-tracking, concepts/ai-history, concepts/biometric-identification, and concepts/inverse-cooking-problem / concepts/inverse-trust-problem (curator-blocked as coined terms since 07-10); person pages entities/anil-seth, entities/orin-kerr, entities/eric-goldman, entities/kevin-klyman, entities/lee-anne-fennell, entities/michael-j-d-vermeer, entities/sacha-altay, entities/joseph-bernstein, entities/roge-karma, entities/alondra-nelson, entities/clayton-christensen, entities/adam-smith, entities/dina-powell-mccormick, entities/raja-krishnamoorthi, entities/samuel-weinbach, entities/ilhan-scheer, entities/david-luan, entities/girish-gupta, entities/alex-mallen, entities/caleb-biddulph; institutions entities/federal-reserve, entities/ecb, entities/bank-of-england (all three from International Monetary Fund (IMF)), entities/carnegie-mellon, entities/santa-fe-institute, entities/sais, entities/uw.
  • Carried from prior runs: slice-11 residue, the standing source-robustness-check queue, none of which moved today (California SB 53 — Transparency in Frontier AI Act in-deg 234, Colorado AI Act (SB 24-205) and SB 25B-004 (Date Amendment) 178, Executive Order 14365 — Ensuring a National Policy Framework for AI 144, Clawed 105, California SB 243 — Companion Chatbots 98); slice-12 residue (NIST AI Risk Management Framework (AI RMF 1.0) in-deg 74, Paris AI Action Summit Declaration (2025), On the Biology of a Large Language Model — routed to source-robustness-check on 07-28); the Anthropic's Responsible Scaling Policy (Version 3.1) status: superseded question left for the curator on 07-27; slice-10 analysis/ residue; slice-9 model/industry residue (DeepSeek-V3, DeepSeek-R1, Qwen3, Kimi K2, Gemma (Google open-weight models), Helix (Figure AI Vision-Language-Action model), LongCat-2.0 (Meituan), IsoDDE (Isomorphic Labs), Defense / Military — AI Deployment, Retail — AI Deployment, Energy and Electric Power Sector, Construction — AI Deployment); industries/manufacturing topical hole; slice-8 residue; H.R. 9619 primary text (congress.gov empty body ×4, not retried); METR "expenditure horizon"; Zitron "Subprime Data Center Crisis"; Pethokoukis transformative-AI essay; the OpenAI/Apollo "metagaming" post and the four arXiv papers surfaced by the 07-26 Redwood pulls; METR's 19 May 2026 frontier risk report; the Zvi Mowshowitz 07-26 follow-up and the Tech Policy Press CFR/Stanford discussion; the college-admissions-essay homogenization study; the Consumer Federation of America / UCLA FTC complaint against Speechify; Substack-redirect URL upgrades.
  • Standing carries (needs-review): claude-code / claude-cowork placement (curator-blocked, 07-09); inverse-cooking / inverse-trust coined terms (07-10); companies/fairly-trained misfile (07-16); entities/cdao / government/cdao duplicate (lint's lane); models/mai-cyber-1-flash quality-gate decline (07-28, with the three triggers that would make it buildable); the 14 dated dashboard-rebuild-failed notes from June, which no run has revisited in six weeks — still worth a curator decision on whether they are live.

One-line summary

Built Compute Thresholds from already-ingested material, rebuilt Safety Cases for Frontier AI from 325 to ~3,400 words against its primary text, folded a third entry-into-force account and the two open EDPB drafts into EU AI Act (Regulation 2024/1689), fixed 4 link aliases and 8 missing sources_count fields; two verified foundational sources on the week's dominant threads — the Pacing the Frontier statement and Hugging Face's forensic timeline — are queued and need the curator's ingest, and the queued Hugging Face task carries a probable mischaracterisation of ExploitGym on two already-ingested pages that should be checked at ingest.