AI Policy Wiki
Dashboard

Topol/Marcus: LLMs and Patient Outcomes — three-document cluster

high confidence · updated 2026-06-06

A 2026-05-03 cluster of three converging documents arguing that current LLMs have not improved clinical patient outcomes — value is concentrated in administrative work. The named 'Topol/Marcus' source is synthesized from Marcus's Substack essay, Topol's separate review, and a Nature Medicine editorial.

"Topol/Marcus: LLMs and Patient Outcomes" is the slug for an argument advanced across three separate documents published within a 48-hour window in early May 2026: a Gary Marcus Substack essay, a longer clinical-side review by Eric Topol, and a Nature Medicine editorial. All three converge on a single empirical claim: as of early May 2026, the published evidence base shows large language models improving administrative healthcare workflows but not yet improving clinical outcomes. The slug names an argument rather than a single co-authored paper.

Provenance

The "Topol/Marcus patient-outcomes gap" is not a single co-authored paper. The argument is synthesized from three separate, real documents that converge on the same conclusion within a 48-hour window:

  1. Gary Marcus, "Have LLMs improved patient outcomes?" — Marcus on AI (Substack), May 3, 2026, https://garymarcus.substack.com/p/have-llms-improved-patient-outcomes . The anchor essay; raw source captured on 2026-05-23 and saved to Raw Sources/Marcus - Have LLMs Improved Patient Outcomes 2026.md.
  2. Eric Topol, "The Paradox of Medical AI Implementation" — Ground Truths (Substack), May 2026, https://open.substack.com/pub/erictopol/p/the-paradox-of-medical-ai-implementation . The detailed clinical-side review Marcus's essay is responding to. Cited via the Marcus essay; full text not yet fetched into Raw Sources/.
  3. Nature Medicine editorial — https://www.nature.com/articles/s41591-026-04389-4 (DOI not yet verified independently of the Marcus essay). The peer-reviewed editorial both pieces cite, and the strongest primary document of the three; to be pulled and verified in a follow-up ingest cycle.

The slug sources/topol-marcus-llm-patient-outcomes is retained as a single hub because the inbound links treat the cluster as one argument. It is referenced from pages including Healthcare — AI Deployment, (Source: anthropic.com), and (Source: openai.com). The page should not be confused with a peer-reviewed paper.

Core claim

Topol's framing in The Paradox of Medical AI Implementation, quoted approvingly by Marcus, is that there is "very little evidence for LLMs benefiting patients or doctors for health outcomes" outside administrative work. All three documents converge on the same empirical claim: as of early May 2026, the published evidence base shows LLMs improving administrative healthcare workflows (note-drafting, ambient scribing, prior authorization, billing, claims appeals) but not yet improving clinical outcomes (diagnostic accuracy that translates to better patient outcomes, treatment-decision quality, mortality, morbidity).

What the Marcus essay says

The Marcus essay frames Topol as someone who "tends to be more bullish" than Marcus on medical AI, so that Topol's negative conclusion carries extra evidentiary weight, characterized as "a believer concedes". Topol's review was, in Marcus's account, "partly inspired by a recent editorial in Nature Medicine" reaching the same conclusion.

Marcus distinguishes this Topol review from his own recent essay "Please don't trust your chatbot for medical advice", which focused on unsupervised patient use; the Topol review covers clinical-outcome studies. Marcus presents the two as converging: both unsupervised lay use and clinician-mediated LLM use lack a positive evidence base on outcomes as of May 2026. His stated caveat is near-term skeptical rather than categorically anti-AI-medicine: "AI will surely someday be a major boon for medicine, but current tools such as domain-general chatbots may not be up to the job."

The Marcus piece runs about 300 words. Its load-bearing work is citation and synthesis rather than independent analysis; the empirical detail sits in the Topol review and the Nature Medicine editorial.

Relation to other wiki coverage

The cluster functions as an empirical constraint on both (Source: openai.com) (ChatGPT Health, Jan 7, 2026) and (Source: anthropic.com) (Claude for Healthcare, April 2026): both vendor launches foreground administrative workflows, where the empirical case is strongest, and leave clinical decision-making to "support, not replace" framing. The cluster also serves as a foil for AI in Healthcare, Healthcare — AI Deployment, and the deskilling counter-evidence in Lancet Endoscopist Deskilling Study (2025).

The argument intersects several other concepts:

  • Enterprise AI Deployment Gap — value capture is concentrated where the evidence is strongest.
  • AI Deskilling, via the Lancet endoscopist study — where LLMs are used clinically, downstream skill erosion may surface in outcomes years later.
  • AI for Science — the clinical-outcomes lag does not contradict basic-science and drug-discovery progress, which are treated separately.

Conditions and scope

The Marcus, Topol, and Nature Medicine convergence is dated May 3, 2026, and is explicitly conditional on the currently published evidence base. The pieces present the conclusion as one that would change if new evidence lands, specifically:

  • A multi-site RCT showing LLM-assisted clinical decision-making improves a hard outcome (mortality, readmission, time-to-diagnosis on a high-volume condition) in clinician-in-the-loop deployment.
  • A registry-level study (for example, national EHR data) showing AI-deployment cohorts outperform controls on a population endpoint over 12 or more months.
  • An FDA-cleared SaMD class beyond the current radiology/cardiology niche showing outcome improvement, not just task accuracy.

Marcus and Topol both anticipate that some outcome evidence will land in 2026–2027. Absent such evidence, the operative empirical state described by the cluster is the administrative-versus-clinical asymmetry.

Relationships