Scanned
Recent window: 4 New Developments Log/ files (2026-08-10 08:11 and 22:10; 2026-08-11 08:12 and 22:05) and 65 wiki pages with last_updated in the last 48 hours. Rotation slice: 13 — sources/ Q–Z (129 pages) plus the Raw Sources/ orphan check.
Broken-link baseline at the start of the run, mainspace only, after normalising the escaped-pipe artifact: 251 distinct broken targets across 279 instances — unchanged from yesterday's figure. The trailing-backslash targets (companies/anthropic\, concepts/vending-bench\, models/gpt-55\) again account for 45 phantom entries across 78 files and are again not a defect: [[page\|Alias]] is correct Obsidian syntax inside a markdown table. This is the twelfth consecutive run to record it; the one-line .rstrip("\\") fix belongs in bin/lint-scan.py, and Wiki/_meta/lint-report.md still publishes the inflated figure of 252.
Two scan components came back clean:
Raw Sources/orphan check (slice 13, second half). All 785 raw files either carry asource_classin frontmatter or are named in the audit log. Zero orphans.- Strict thin-anchor rule in slice 13 (in-degree ≥6 and under 350 words, in-degree computed per qualified target rather than by basename). Zero hits, the third consecutive slice with none.
sources/the-industrial-explosionis the nearest miss at 358 words and in-degree 10, and isconfidence: high. Thirteen slice-13 pages trip the weakermedium-with-sources_count: 1test; they aresource-robustness-checkwork, not gap-identifier expansion, and are listed in the deferred backlog.
Because the slice produced no expansion candidate, the thin-anchor scan was widened wiki-wide under the same strict rule. That returned seven pages, one of which — concepts/ai-and-misinformation — was directly implied by the recent window and was actioned.
The open queue was empty of INGEST- tasks at the start of the run: all three tasks queued on 2026-08-11 (ARI, Zuckerberg, IFP) plus the APA advisory queued on 2026-08-10 have been ingested, and their sources/ pages appear in the 48-hour edit window.
Twelve candidates survived alias resolution and deduplication against the open queue, the seven gap-scan reports since 2026-08-05, and existing Raw Sources/ files. Seven were actioned, at the low end of the 6–10 cap; nine web pulls were made.
Gaps actioned (7 of 12 found)
New pages created (live)
- EU Code of Practice on Transparency of AI-generated Content — the EU Code of Practice on Transparency of AI-generated Content. A named EU instrument with its own signatory list, its own effective date and its own drafting history, and the direct object of a last-48h development (Anthropic's 11 August watermarking commitment), yet present in mainspace only as prose inside EU AI Act (Regulation 2024/1689) and one line on Anthropic — while its sibling instrument, the GPAI Code of Practice, has both a
legislation/and asources/page. Built from nine sources, four of them the Commission's own canonical pages: the policy page, the signatory-list news article, the signing FAQ, and the June press release, plus Cooley and Tech Policy Press on the compliance posture. Records the full drafting timeline (September 2025 consultation through the 10 June 2026 closing plenary), the two-section structure, the working-group mandates, the adequacy assessment, the Signatory Taskforces due in September 2026, and the signatory counts.confidence: medium. Nine inbound links added. - NVIDIA Nemotron — NVIDIA's open-model family. Named on ten mainspace pages and used as a benchmark comparison baseline on four of them (Inkling, Inkling-Small, Inkling Model Card (Thinking Machines Lab, July 2026), Open-Source AI / Open-Weight Models) with no page behind it, and reinforced twice in the recent window: the Nemotron 3.5 Lightning release of 11 August 2026 and the reported Nemotron 4 family. Built from seventeen sources, six of them NVIDIA's own domains (research.nvidia.com, nvidianews, developer.nvidia.com, build.nvidia.com, the NVIDIA-NeMo GitHub repository), plus CNBC and Ollama on the 3.5 Lightning release. Covers the Nano/Super/Ultra configuration, the hybrid Mamba-Transformer MoE architecture, LatentMoE, MTP and NVFP4, the published training datasets, and NVIDIA's own throughput comparisons.
confidence: medium. Six inbound links added.
Pages expanded (live)
- AI and Misinformation — the strongest thin anchor wiki-wide: in-degree 11, 338 words,
confidence: medium,sources_count: 0, and by its own text a placeholder ("This page serves as a thin umbrella anchoring 4+ wiki references"). Expanded to 1,630 words from five supporting web sources, and the four house-style violations removed (three instances of wiki self-reference plus the "thin umbrella" line in the description). Added: a measured-prevalence section with a five-row table of 2024-cycle estimates (Meta under 1% of fact-checked misinformation; Logically Facts 1.35% of 1,695 fact-checks; India fewer than two dozen of 1,858 viral WhatsApp messages; Bangladesh 1.9%; US about 6%); the Alan Turing Institute counts (16 confirmed viral cases in the UK general election, 11 across the EU and French elections) and its CETaS finding of no measurable effect on results in the UK, EU, Taiwan and India; the CISA and ODNI 2024 assessments verbatim; the Hackenburg–Margetts and Jones–Bergen persuasion findings; the countervailing PNAS result that all four types of persuasive AI tested produced significant attitude change; the supply-versus-demand argument; the cheap-fakes-versus-deepfakes finding; and model output as an error channel. Every pre-existing fact, wikilink and section survives.confidenceheld atmedium;sources_count0 → 6. Original backed up toWiki/_meta/_revision-backups/concepts/ai-and-misinformation-2026-08-12.md.
Link and accuracy fixes
No alias errors were found; the 45 apparent trailing-backslash targets are correct table syntax (see Scanned). Fifteen edits were made across eleven pages, connecting the two new pages into the graph and correcting one substantive mis-citation:
- Data Provenance, C2PA, and Watermarking — mis-citation corrected. The page stated that EU AI Act Article 50's machine-readable marking requirement is operationalised by the GPAI Code of Practice. It is not: the GPAI Code covers Articles 53 and 55, and Article 50 is operationalised by the Transparency Code. Corrected, with both instruments now named and distinguished. This is a wrong attribution of a legal instrument, not a formatting issue.
- Anthropic — "the EU AI Act's Transparency Code, which took effect on August 2, 2026" recast so that the date attaches to the Article 50 obligations rather than to the voluntary code, and wikilinked to the new page. The code was published 10 June 2026; 2 August 2026 is the date the underlying legal obligations became applicable.
- Nvidia & TSMC — AI Compute Infrastructure — Nemotron 3 Ultra and Nemotron 4 wikilinked; NVIDIA's own release date (4 June 2026) and configuration (550B total / 55B active) added alongside the newsletter-sourced "around June 2, 2026"; the Nemotron 3.5 Lightning release folded in.
- EU AI Act (Regulation 2024/1689) (2 edits), EU General-Purpose AI Code of Practice (2025), AI Content Provenance, Synthetic Media / Deepfakes (2 edits, including a typed
related:relationship) — Transparency Code references wikilinked; the Commission's own 5 August 2026 signatory counts (about 190 organisations; Section 1: 82, Section 2: 152) added where the pages carried only the Guardian's "more than 180." - Gemma (Google open-weight models), Inkling, Inkling-Small, Open-Source AI / Open-Weight Models, Inkling Model Card (Thinking Machines Lab, July 2026) — plain-text Nemotron mentions wikilinked.
No sourced fact was dropped and no unsourced claim added in any edit.
Queued — foundational sources
- Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping and Maksym Andriushchenko, "Stealing Reasoning Traces from Proprietary LLM APIs" (arXiv:2608.09867, 10 August 2026) — the highest-scoring gap of the run (8). A dangling foundational reference: the 2026-08-11 evening digest names the finding, quotes Panfilov and reports Anthropic's response, but the paper has no
sources/page and noRaw Sources/file. Verified:arxiv.org(canonical host for a preprint), fetched live, HTTP 200; identity corroborated across arXiv's own citation metadata, the Wired coverage and the submission-history line; identifier check passes on arXiv:2608.09867v1 and the DataCite DOI 10.48550/arXiv.2608.09867. Licence CC BY 4.0. Queued:INGEST-stealing-reasoning-traces-panfilov.md. Verification record:Wiki/_meta/queue/gap-scan/proposed-sources/stealing-reasoning-traces-panfilov-2026.md.
No raw file saved, by design. The HTML render is about 490,000 characters and the PDF about 16 MB. Protocol step 5 rejects saving a summary or excerpt in place of full text — the exact defect found in the 2026-04-14 APA compilation on 2026-08-10 — so the task carries both verified full-text URLs for a fresh fetch at ingest rather than a partial capture.
Source-fidelity findings
Four discrepancies between digest text and primary sources. Two were corrected in mainspace this run; two are recorded in the ingest task for resolution against the primary text.
- A voluntary instrument's date conflated with a statutory one. The 2026-08-11 morning digest, and Anthropic following it, describe "the European Union AI Act's Transparency Code, which took effect on August 2, 2026." The code is voluntary and was published 10 June 2026; what took effect on 2 August 2026 are the Article 50 obligations it operationalises. Corrected on Anthropic and stated correctly on the new page. This is schema v4.7 failure class 2 — compression that reverses a claim's modality, here treating a voluntary code as though it were the binding instrument.
- A legal instrument mis-attributed. Data Provenance, C2PA, and Watermarking attributed the Article 50 marking requirement to the GPAI Code of Practice. Corrected; both codes now named and distinguished.
- A restatement's date displacing the document's. Three mainspace pages date the Nemotron 3 Ultra release to "around June 2, 2026," from a newsletter. NVIDIA's own research site dates it 4 June 2026. Both now recorded on Nvidia & TSMC — AI Compute Infrastructure and the new model page, with the primary date given precedence.
- Unverified specifics carried from secondary coverage. The 2026-08-11 evening digest reports, from Wired, that the Panfilov team fed 90 questions to each model, that Kimi K3's output was "strikingly similar" to hidden traces of Claude Opus 4.8 and GPT 5.6 Sol while DeepSeek and Inkling's were not, and that the paper states the work "cannot causally establish distillation." None of this appears in the abstract. The affiliations Wired gives (Tübingen, Max Planck, MATS, Snyk) are likewise not on arXiv's author block. All flagged in
INGEST-stealing-reasoning-traces-panfilov.mdnotes 4, 6 and 7 for confirmation against the paper body before folding. The distillation caveat in particular materially limits the claim and must not be dropped.
Authenticity-verification failures
None. One source was pulled through the full protocol and passed every step; the eight other pulls were supporting web sources folded as (Source: URL) and not subject to the protocol.
The Panfilov paper is the first source in six runs to pass the identifier check on a formal identifier rather than on a canonical URL alone — it carries both an arXiv ID and an arXiv-issued DataCite DOI. The standing weakness recorded on 2026-08-08, 2026-08-10 and 2026-08-11 (think-tank reports, corporate newsroom posts and agency advisories carrying no DOI, ISBN or docket number) does not apply here.
One limitation is recorded rather than waved through: the paper's institutional affiliations are corroborated only by Wired at this stage, since arXiv lists the eight authors without affiliations. No entity page or affiliation claim should be written from that attribution until the paper's author block is read.
Deferred backlog (over the daily cap — re-surfaces next run)
The first five entries are the deferred scored candidates making up the 12-found / 7-actioned arithmetic. The rest are standing carries and resolved non-gaps, recorded so they stop re-scoring.
companies/riot-platforms(score 3) — the Bitcoin miner that disclosed a 20-year, 191 MW supply agreement on 10 August 2026 that Bloomberg identified as a $9.1 billion contract with Anthropic. Zero prior mainspace mentions. A public company with a substantial record, so the quality gate is passable, but a single Bloomberg report supplies only the contract terms; the facts belong in Anthropic and AI Data Centers via the nightly fold until there is enough for a page. Same disposition ascompanies/theseus-infrastructureyesterday, which also remains open.- CoreWeave staleness (score 3) —
last_updated: 2026-07-23, in-degree 16, and materially behind on fast-decay financials after the 11 August Q2 report (revenue $2.58B, up 112% year over year; net loss $626M; $104B backlog; 1.5 GW active power; $35B debt; FY capex guidance raised to $35–39B). Not a thin anchor (839 words,confidence: high,sources_count: 6), so outside the AUTO-expand lane, and folding a dev-log item isdevelopments-log's job, not this skill's. Flagged for the nightly fold. - Six concept pages forward-referenced as "(planned)" (score 1 each) —
concepts/algorithmic-decisionmaking(from Toward Responsible AI in Health Insurance Decision-Making — Mello, Trotsyuk, Djiberou Mahamadou, Char (Stanford HAI Policy Brief, February 2026), including a typedsupports:relationship),concepts/talent-flow-china-us(from The AI Grand Bargain — Ben Buchanan and Tantum Collins (Foreign Affairs, October 2025)),concepts/ai-and-language-models(from Fair Learning — Mark A. Lemley & Bryan Casey (Texas Law Review, 2021)),concepts/economic-possibilities-for-artificial-intelligence(from How I Choose Which Cloudflare Employees to Replace With AI),concepts/a-vision-of-democratic-ai(from AI and Democracy),concepts/data-broker-regulation(from Connecticut SB 4 — Data-broker registration + geolocation-sales ban + facial recognition (CTDPA amendment)). Each has exactly two inbound references from a single page, and several are explicitly marked "(planned)" by the ingesting run — deliberate forward references rather than accidental breakage. Worth one dedicated pass rather than one page per day. - Author entity pages implied by
sources/pages, none created (score 1 each) — Lee Anne Fennell, Kevin Klyman, Orin Kerr, Eric Goldman, Michael J. D. Vermeer, Alondra Nelson, Sacha Altay, Joseph Bernstein. Each is two broken inbound links from its own source page. A batch pass, not a daily gap. - Near-threshold thin anchors in slice 13 (score 2) — thirteen
sources/Q–Z pages at in-degree ≥8 withconfidence: mediumandsources_count: 1, none under the 350-word bar:stanford-hai-ai-index-2026(in-degree 35),science-of-scheming(26),safety-cases-frontier-ai(21),taxonomy-systemic-risks-gpai(15),redwood-mallen-notes-evade-containment-2026(11),trustworthy-agents-in-practice(10), and seven others at 8.science-of-schemingis the strongest candidate at 739 words for a page with in-degree 26. All aresource-robustness-checkwork. - Resolved as non-gaps this run, recorded so they stop re-scoring — Facial recognition: the Western Australia Police live-FR trial (11 August digest) suggested a
concepts/facial-recognitiongap, but AI and Surveillance carries a substantial## Facial recognitionsection covering Clearview, NIST FRVT/FRTE, public-housing deployment and the wrongful-arrest litigation, and Surveillance Technology covers the adjacent video-analytics terrain. Not a gap; the WA trial is a fold for AI and Surveillance. Chipflation: the Morgan Stanley term from the 11 August morning digest is already folded at Semiconductor Supply Chain, and coining a page around a bank's neologism would breach house style §2c. Not a gap. Genome language models: the King–Hie Science paper of 6 August already has Generative design of bacteriophages with genome language models. Dropped on dedupe, correctly. - Persons named in this window with no page — Greg Casar and Doris Matsui (leads of the 29-Democrat letter to OpenAI); Michael Dalton (OpenAI researcher, Black Hat disclosure); Alexander Panfilov and seven co-authors; Michael Aciman (Anthropic spokesperson); Mike Intrator (CoreWeave CEO); Col Blanch, Peter Collins and Malcolm Crompton (WA facial-recognition trial). All at one mention. Casar and Matsui are the strongest and will recur if the requested hearings happen.
- Organisations named in this window with no page — Riot Platforms, Snyk, MATS Research, NEC, Trajectory Labs, CivAI, the AI Psychological Research Coalition, SpecterOps, Theseus Infrastructure, Macquarie Asset Management, GIC. Standing carries also still below threshold:
entities/steven-adler,entities/consumer-technology-association,entities/epic,entities/cpsc,entities/national-academies,entities/seth-lazar,entities/vector-institute, MeitY. sources/protecting-wellbeing-of-usersandsources/when-the-interface-is-neural— both resolved and marked superseded on 2026-08-10; deletion still needs the curator's sign-off atWiki/_meta/queue/gap-scan/needs-review/2026-08-10-two-stub-sources-pages-superseded.md. Third carry.- Ambiguous bare wikilinks — the systemic finding recorded on 2026-08-10 (roughly 1,700 link instances whose bare basename exists in two folders, led by
[[eu-ai-act]]at 250 and[[americas-ai-action-plan]]at 211) is unchanged and remains scopedlintwork. - Recent-window items the nightly fold already handled, not gaps — the Nvidia $500B financing alliance and the 25% backstop option, Intel's upsized $20B offering, OpenAI's $7B tender, Anthropic's IPO preparation, the OpenAI agent message-board disclosure, Manus's unwinding from Meta, Muse Spark 1.2 open-weighting, Caitlin Kalinowski's move.
One-line summary
Twelve gaps found, seven actioned: two pages created live (EU Code of Practice on Transparency of AI-generated Content, NVIDIA Nemotron), the wiki's strongest thin anchor AI and Misinformation expanded from 338 to 1,630 words with six sources where it had none, fifteen link-and-accuracy edits across eleven pages including a corrected mis-attribution of the EU Article 50 marking requirement to the wrong code of practice, and the Panfilov reasoning-trace-extraction paper verified on arXiv and queued. A verification pass was run over all three pages against the house-style checklist before the run closed; it found and fixed two coined/unattributed superlatives, one added wiki self-reference, two mis-typed regulated-by: relationships, two understated sources_count values, a reconciliation gap between NVIDIA's announcement and released parameter counts, and one dropped candidate-source name pair (Stanford Internet Observatory, RAND) in the expanded page. One citation was dropped rather than kept: an AWS SageMaker JumpStart URL whose date path (/2026/01/) is inconsistent with the 11 August 2026 release it was cited for.
Needs the user's review: the single INGEST- task, and specifically whether the Wired-sourced specifics in the 2026-08-11 evening digest (the 90-question protocol, the Kimi K3 similarity finding, the institutional affiliations) survive contact with the paper body — the digest currently carries them as established, and the abstract supports none of them.