AI Policy Wiki
Dashboard

Gap Scan — 2026-08-22

Daily gap hunt — what the wiki is missing and how each gap was triaged.

Scanned

Recent window: 3 New Developments Log/ files (2026-08-20 22:09, 2026-08-21 08:10, 2026-08-21 22:05), 102 wiki pages with mtime inside 48 hours. Rotation slice: 9 — models/ + industries/ (81 pages).

bin/lint-scan.py run at open: 235 distinct broken targets across 259 refs, 24 referenced 2+ times. bin/primary-doc-scan.py --days 3: 6 digests, 3 cited primary-document URLs of which 2 have no raw file, 0 possible matches, 7 items in the keyword pre-filter lane.

One digest is unprocessed. 2026-08-21-2205-ai-developments.md has no matching Operation: Developments-Log entry. The skill's run order assumes gap-identifier follows the nightly fold, so items from that file were treated as genuine gaps rather than as work the fold would have taken; three of them are actioned below. If the fold runs later it should check for overlap.

Gaps actioned (8 of 31 found)

New pages created (live)

  • United States v. Ding — score 8 (live thread +3, dangling foundational reference +3, national-security core +2). "Linwei Ding" appears on 12 content pages including Overview, Export Controls (AI), Chip Smuggling and Export-Control Evasion and Taiwan v. Chen Li-ming et al. (TSMC 2nm Trade-Secret Case), and a sources/ page for the DOJ conviction release has existed since June, but the case had no litigation/ case-tracking page — the folder the "Where does X go?" rules assign to an individual court case. Alias resolution found nothing: litigation/ holds hachette-v-google, lg-munich-google-ai-overviews and us-v-google-search, none of them this matter. Built from Courthouse News's report on the August 20 order, the two DOJ press releases (original March 2024 indictment and February 2025 superseding indictment), Reuters via BusinessWorld, and The Next Web. 1,562 words, confidence: medium, 6 sources.

The page corrects the existing record in two places. First, DOJ: Former Google Engineer Convicted of Economic Espionage (Linwei Ding) records fourteen standing convictions; seven no longer stand, and the new page declares contradicts: against it rather than silently diverging. Second, the digest describes the judge as having "threw out seven economic espionage counts"; the order is a judgment of acquittal under the reasonable-doubt standard, which the court noted bars retrial on double-jeopardy grounds — a materially stronger disposition than "thrown out" conveys.

  • Cloverleaf Infrastructure — score 5 (live thread +3, compute/infrastructure core +2). Zero occurrences anywhere in the content folders before today. Nvidia's minority investment was announced August 21 and is the third in a short sequence alongside SB Energy and Lancium, both of which have pages; Cloverleaf did not. Built from Reuters, The Information, the company's own newsroom, the New York Times, Latitude Media and the WSJ notice of the 2024 founding round. 1,031 words, confidence: medium, 6 sources.

Source-fidelity correction carried onto the page. The 2026-08-21 evening digest calls Brian Janous "Cloverleaf chief executive." Cloverleaf's own site consistently styles him co-founder and chief commercial officer; a search snippet elsewhere gives "chief strategy officer." The company's own designation is used. The digest also names only SB Energy as the precedent investment; The Information names Lancium as well.

The page records the 10× and 90% energy-cost figures as the company's own projections rather than measured results, and notes under ## Open questions that the published figures do not appear to net out launch, replacement and radiation-degradation costs. TechCrunch's own reporting that the economics of orbital AI are difficult is cited alongside the company's claims.

Pages expanded (live)

  • Defense / Military — AI Deployment — score 6 (slice-9 thin anchor at in-degree 18 with 605 words and last_updated: 2026-07-13, +2 in-degree, +2 national-security core, +2 live thread via the NSSTS). Was carrying the sector's frontier-model procurement story only through the Anthropic-Pentagon dispute, with no mention of either procurement wave. Added: the July 2025 CDAO awards of up to $200M each to Anthropic, OpenAI, Google and xAI; the May 1, 2026 classified-network agreements with eight companies, sourced to the War Department's own release (IL6/IL7 environments, the AI Acceleration Strategy's three tenets, the stated anti-vendor-lock architecture, and the GenAI.mil figures of 1.3M+ users and hundreds of thousands of agents in five months, flagged as the department's own unverified numbers); Anthropic's absence from that list and DefenseScoop's account of why; Cameron Stanley as CDAO; and a new strategy section built on the NSSTS. 605 → 1,244 words, 8 → 12 sources, confidence: high retained and now better supported. Backup at Wiki/_meta/_revision-backups/industries/defense.md.
  • Gemma (Google open-weight models) — score 4 (slice-9 thin anchor, in-degree 8, 318 words, +1 in-degree, +2 open-weight policy core, +1 backlog). The page's own text flagged an unresolved discrepancy about the Gemma 4 release, and the discrepancy turned out to be an error rather than an ambiguity. Two newsletter items dated Gemma 4 to June 2 and to the week ending July 11, 2026, and named a "12B" variant. Google's announcement and Ars Technica's contemporaneous report both place the launch at April 2, 2026, with four sizes — 26B MoE (3.8B active), 31B dense, E4B and E2B — and no 12B among them. Added: the four-size lineup and its deployment split, context windows (256k large / 128k edge), 140+ languages, the Arena debut claim, the Apache 2.0 licence switch and the specific developer objections to the prior Gemma Terms of Use, and the Gemini Nano derivation chain (Nano 3 from Gemma 3n; Nano 4's 2B/4B from Gemma 4 E2B/E4B). 255 → 1,231 words (against the retained older backup), 4 → 8 sources, confidence: medium held.

Both newsletter claims and their citations were preserved, in a subsection that states the conflict and says which side is treated as established and why, rather than being deleted as wrong.

None. Consistent with the settled finding recorded across the last five runs: the remaining broken targets are author surnames, institution shorthand and deliberate forward references. The referenced-2+ list was not re-audited.

Queued — foundational sources

  • National Fair Housing Alliance, comment to the House Financial Services Committee on the AI Risks and Modernization RFI (2026-08-14) — gap type 9, cited lane. Named in the 2026-08-21 morning digest's ## Primary documents block on an Inside AI Policy report whose body was unretrievable. Verified: nationalfairhousing.org, the organization's own domain, 12 of 12 pages. Saved: Raw Sources/NFHA - Comment on HFSC Democrats RFI on AI Risks and Modernization in Financial Services (2026-08-14).md. Queued: INGEST-nfha-hfsc-ai-rfi-comment.md. Verification record in proposed-sources/.

The capture changes what the source supports. The coverage carried one proposition — no safe harbour for any AI actor. The letter answers five numbered RFI questions, four of them absent from the coverage entirely, and states the liability position more precisely: liability allocated by role across developers, deployers, third-party vendors and users, which is not the same claim as uniform liability. It is also addressed to Ranking Member Maxine Waters personally, not to "Democrats on the committee" as the digest has it.

  • Office of Science and Technology Policy, "National Security Science and Technology Strategy" (August 2026) — gap type 9, cited lane. Verified: whitehouse.gov, HTTP 200, 24 pages. Queued URL-only under the large-PDF allowance: INGEST-nssts-2026.md. A 12-page verification read was taken and its findings recorded in the ingest task; that read is explicitly not a capture, and the task flags that Appendix A — an update to OSTP's list of critical and emerging technologies — falls outside it and must be present in the eventual capture.

A correction is attached for the ingest: the digest's phrases "sensor fusion", "foundation models" and "autonomous command and control" do not appear in the strategy's own text, which says C5ISR, "future computing technologies" and "AI and autonomy". The gloss came from the trade coverage.

  • NIST SP 1353 (Initial Public Draft), CSF 2.0 quick-start guide for using AI for CSF analysis and reporting — gap type 9, named lane, promoted after reading the item. Verified: csrc.nist.gov, DOI 10.6028/NIST.SP.1353.ipd, document history 08/19/26. Queued URL-only: INGEST-nist-sp-1353-ipd.md, because the guide's substance sits partly in four supplemental ZIP archives carrying the prompts and the simulated corporate documents, and a PDF-only capture would miss them.

Send-date-as-event-date error found and flagged. The digest records the release as becoming public on August 21, 2026. NIST's own record, its citation_publication_date metadata and its document-history line all give August 19, 2026. This is the exact failure class .claude/skills/source-fidelity/SKILL.md exists to catch; it has propagated to National Institute of Standards and Technology (NIST) and should be corrected there.

Authenticity-verification failures

None. Three documents pulled, three verified against canonical hosts.

One document was not pulled on protocol grounds: Judge Chhabria's 18-page order granting the judgment of acquittal in United States v. Ding. The only reachable copy is a Courthouse News reproduction at courthousenews.com/wp-content/uploads/2026/08/usa-v-ding-granting-motion-for-judgment-of-acquittal-economic-espionage.pdf. The protocol prefers the court's own docket or CourtListener for court filings; the CourtListener MCP required authentication unavailable in this session. Deferred as a primary-document candidate rather than saved from a secondary host.

Found while link-checking the Defense / Military — AI Deployment revision, then widened. Fifty-three page basenames occur in more than one content folder; 41 are the target of bare [[wikilink]] references; those references number 2,442. Each resolves to an undetermined one of two pages. For scale, the broken-link count lint reports today is 259 references — the ambiguous count is roughly ten times larger and is measured nowhere.

Most are the sanctioned sources/X + legislation/X pairing with an unsanctioned shared slug ([[eu-ai-act]] 250 refs, [[americas-ai-action-plan]] 211, [[california-sb-53]] 160, [[colorado-ai-act]] 141, [[eo-14365]] 132, [[nist-ai-rmf]] 94). Two are genuine "No duplication rule" violations, the same organization in two folders: government/cdao / entities/cdao, and government/cisa / entities/cisa.

bin/lint-scan.py cannot see this: it asks whether a target exists, and these exist twice. Nothing was changed — neither remedy is mechanical, and a bulk retarget that guesses wrong on a few hundred of 2,442 references would encode the wrong answer rather than leave the question open. Full write-up and suggested disposition in Wiki/_meta/queue/gap-scan/needs-review/2026-08-22-ambiguous-bare-wikilinks.md.

Deferred backlog (over the daily cap — re-surfaces next run)

  • models/deepseek-v3 (score 4), routed rather than deferred. In-degree 49 against 608 words — the most-relied-upon underbuilt page in slice 9. Not expanded here because the work it needs is benchmark tables, lineage and peer comparison for a December 2024 model, which is model-robustness-check's lane, not a supporting-source top-up. Route it there rather than re-scoring it as a thin anchor each fortnight.
  • Remaining slice-9 thin anchors (score 2–3 each). models/helix-figure (in-degree 6, 208 words, 1 source), models/alphagenome (6, 223 words, high on 2 sources — the thinnest high-confidence page in the slice), models/longcat-2 (6, 317 words). industries/consulting (in-degree 5, 214 words, low/1), industries/accounting (3, 197, low/1) and industries/construction (1, 152, low/1) are thin but below the in-degree band.
  • Live-thread items named in the window with no page (score 3 each). Vercel (13 content mentions, no company page; the $1M HackerOne sandbox-escape programme closes September 1, 2026, so it decays); Slack (Slack Code launch, agent channels); the DOJ–TikTok $400M COPPA settlement, which has no litigation/ page though Children's Online Privacy Protection Act (COPPA) was created on 2026-08-20; H.R. 10110 Housing Price Transparency Act (algorithmic rent-setting, adjacent to Surveillance Pricing and the 16 pages mentioning RealPage); Representative Barrett's two data-center bills of August 20; Mistral Agentic Search; Micron's $10B Boise lab (page exists at Micron Technology, a fold not a gap); Broadcom's $60B+ AI debt raise; Cloverleaf's peer Manhattan West Ventures.
  • Private Safety Processing (score 5), third consecutive deferral. Carried in three digests now (August 19 evening, August 20 evening, August 21 evening) and folded into OpenAI and Anthropic, with Anthropic's parallel data-retention change alongside it. The 2026-08-20 report judged it "likely folds into an existing privacy or safety concept page rather than warranting its own." That judgment has now been made three times without being acted on either way. Per the three-run rule this is structural: the question is whether the wiki should hold a concept page on privacy-preserving safety monitoring, and it should be decided rather than re-deferred. OpenAI's technical white paper is due September 2026 and would be the natural anchor.
  • Primary-document backlog carried from the 2026-08-20 baseline (score 5–6). The White House release of 2026-07 (instrument still unidentified), the Alphabet Q2 2026 earnings release, and the ANI v. OpenAI Delhi High Court order of 2026-07-24 where the cited host is a blog reproduction. Add the Chhabria acquittal order, above.
  • The 24 referenced-2+ broken targets, unchanged — eighth consecutive run. Twelve are author entity pages implied by sources/ pages, each at in-degree 2. Recorded once, not re-listed: a batch pass over the author-page set is the right instrument and has been the stated finding since 2026-08-15.
  • Institution shorthand targets (score 1–2), unchanged, seventh run. entities/carnegie-mellon, entities/santa-fe-institute, entities/bank-of-england, entities/ecb, entities/federal-reserve, entities/uw.
  • Concept targets named but pageless (score 2 each), unchanged, eighth run. concepts/inverse-cooking-problem and concepts/inverse-trust-problem (parked under needs-review/2026-07-10), concepts/a-vision-of-democratic-ai, concepts/data-broker-regulation, concepts/talent-flow-china-us, concepts/economic-possibilities-for-artificial-intelligence, concepts/ai-and-language-models, concepts/algorithmic-decisionmaking (disposition note at needs-review/2026-08-19).
  • The 29 date-stamped headers in lint-report.md (score 1). Unchanged for six cycles; the report still does not distinguish legitimate procedural histories from timeline dumps, which remains the finding.

Checks run before closing

  • No broken outbound links introduced. All wikilinks on the three new pages and two expanded pages were resolved against the page set after writing. Three were caught and fixed during the run: [[companies/spacex]] (does not exist; the page's placement is parked under needs-review/2026-08-18) was rewritten to plain text, a placeholder [[sources/…]] written during drafting was rewritten as prose, and [[dod-directive-3000-09]] was found to be ambiguous rather than broken and was left as written — see the structural finding above.
  • Broken-link count unchanged at 235 distinct / 259 refs. Expected, and the same pattern as the last three runs: none of the actioned gaps was on the broken-link list, because all were dangling prose references, live-thread items or primary documents.
  • Style scan clean. All five touched content pages checked programmatically against the banned-header, hyperbole, self-reference, standing-Predictions and date-header lists; zero hits on each. Wiki-wide counters after the run are unchanged: banned headers 0, self-references 0, standing Predictions 0, header hyperbole 0, date headers 29.
  • Revision fidelity verified mechanically. URL and wikilink sets diffed old against new for both revisions: zero URLs dropped, zero wikilinks dropped. industries/defense 605 → 1,244 words; models/gemma 255 → 1,231 against the retained backup.
  • A backup was missed and is recorded rather than papered over. Wiki/_meta/_revision-backups/models/deepseek-v3.md was taken at the start of the run on the assumption that deepseek-v3 would be the second expansion; when the choice changed to models/gemma, no fresh gemma backup was taken before the rewrite. An older gemma backup existed and was used for the mechanical diff, and the pre-revision text was additionally checked fact-by-fact from the read taken earlier in the run — all eleven of its concrete claims and all six of its citations survive. The gap in procedure is real and is noted so the next run does not assume the backup is contemporaneous.
  • Confidence set honestly. All three new pages are confidence: medium; none reached high, and the three-independent-source bar was not met for any. industries/defense retains the high it already carried, now on 12 sources rather than 8. models/gemma held at medium despite reaching 8 sources, because the Arena ranking and the Nano 4 roadmap are fast-decay claims resting on a single April 2026 report.
  • Cap respected. Eight gaps actioned against a 6–10 cap; thirty-one candidates found. Eleven web pulls and four searches, no crawls.

One-line summary

Three new pages (United States v. Ding, Cloverleaf Infrastructure, Starcloud), two slice-9 anchors expanded (Defense / Military — AI Deployment, Gemma (Google open-weight models)), three verified primary documents queued for review including the NFHA comment captured in full; three source-fidelity errors corrected in passing, and a previously unmeasured class of link defect — 2,442 ambiguous bare wikilinks across 41 duplicated basenames — surfaced for the curator's decision.