AI Policy Wiki
Dashboard

Healthcare — AI Deployment

high confidence · updated 2026-07-26

AI adoption in healthcare: ambient scribes, CADe/CADx, drug discovery, claims processing, AI-assisted diagnosis. Key deployers: Novo Nordisk, Epic, hospital systems. Key tensions: deskilling, FDA SaMD regulation, hallucination risk.

AI adoption in healthcare spans clinical (diagnosis, scribes, imaging), operational (scheduling, billing, claims), pharmaceutical (drug discovery, clinical trial optimization), and patient-facing (chatbot triage, mental-health apps) applications. Reviews and benchmarks through 2026 concentrate documented value in administrative workflows while clinical decision-making remains positioned as decision support rather than replacement.

Adoption patterns

Physician use of AI rose sharply across 2023–2026. Per AMA Physician AI Sentiment Report (2026), reported AI adoption among physicians roughly doubled from 38% to 81% since 2023; 70% view AI as a tool against burnout, while 88% expressed concern about skill loss.

Medical-search and reference tools reached broad use among practising physicians. OpenEvidence was used by roughly 65% of US doctors across about 27 million clinical encounters in April 2026, with roughly 650,000 active US physicians and 1.2 million international users; 60% of all queries concerned clinical decision-making (NBC News, May 13). CEO Daniel Nadler said, "We did the hardest thing in the history of American health care…We got the majority of American doctors to all voluntarily adopt a single technology platform." The tool is free for credentialed users and ad-supported. Reported concerns include hallucinations, the lack of rigorous patient-impact studies, and deskilling; MaineHealth asks clinicians to refrain from entering protected health information (PHI). OpenEvidence raised $700M at a $12B valuation (Sequoia, GV, Nvidia, a16z, Thrive), a roughly 12× increase in valuation in just over a year. (Source: nbcnews.com) By July 2026 OpenEvidence was reported to be considering a $200 million raise at a roughly $20 billion valuation, with annualized revenue nearing $300 million after doubling in seven months, 90% gross margins, and separate acquisition talks with a large technology company (Source: theinformation.com).

By application area, ambient clinical documentation, computer-aided detection and diagnosis (CADe/CADx) in radiology, pathology, and endoscopy, and pharmaceutical drug discovery represent the most developed deployment categories. The US Food and Drug Administration (FDA) had cleared 500+ AI/ML-enabled Software as a Medical Device (SaMD) products as of 2025.

Representative deployments

Provider and payer systems. Novo Nordisk, the maker of Ozempic, uses AI for document generation of sensitive pharmaceutical documents (The Information: "Ozempic Maker Says AI Is Finally Reliable Enough to Produce Sensitive Documents"). Epic Systems integrates ambient clinical scribes into its EHR (industry reporting). Multiple hospital systems deploy CADe/CADx in radiology, pathology, and endoscopy (Lancet Endoscopist Deskilling Study (2025)). UnitedHealth Group disclosed approximately $1.5B in AI investment in its Q1 2026 earnings (April 25, 2026), against $111.72B in quarterly revenue; the company describes AI as a pillar of CEO Stephen Hemsley's turnaround, and the figure is the largest declared AI capital expenditure by a US health insurer/services group (Source: Fortune, April 25, 2026).

Health-system-owned and collaborative models. Microsoft paired its seven new MAI models (Microsoft) with a Mayo Clinic collaboration to build a Mayo-owned frontier healthcare model, an instance of a major health system commissioning a proprietary, institution-owned frontier model rather than licensing a general one (June 2, 2026). (Source: microsoft.ai)

Diagnostic and imaging research. A radiomics model described in Gut / BMJ (April 22, 2026), referred to as REDMOD, detects pancreatic cancer signs on routine CT scans before radiologists; it was validated on 2,000 abdominal CTs (73% detection vs. 39% for radiologists, with a 475-day prediction lead time). See FDA — Food and Drug Administration (AI Deployer). (Source: gut.bmj.com) Google DeepMind's AI Co-Clinician (May 1, 2026), a multimodal research effort pairing Gemini with Project Astra to handle telehealth calls and guide live physical exams, matched primary-care doctors on 68 of 140 consultation skills in tests. See Google DeepMind. (Source: deepmind.google) OpenAI and Boston Children's Hospital reported on June 18, 2026, in a paper published in NEJM AI, that OpenAI's o3 Deep Research model helped clarify 18 diagnoses among 376 children with previously unsolved rare diseases — a roughly 5% yield. Researchers at the hospital's Manton Center fed the model clinicians' notes, symptom descriptions, and filtered gene lists for cases spanning neurodevelopmental, neuromuscular, sudden-death, and early-childhood-psychosis conditions, with human geneticists reviewing every output before any diagnosis. Lead researcher Catherine Brownstein called the result "a total game changer," while she, OpenAI health lead Ashley Alexander, and outside experts cautioned against overhyping the tools and stressed that results still require rigorous human review; OpenAI helped fund the project (Source: nbcnews.com). See OpenAI.

Drug discovery and AI-designed biologics. Google DeepMind's drug-discovery spinoff Isomorphic Labs runs a proprietary IsoDDE engine built on AlphaFold 3, with an oncology and immunology pipeline, Lilly and Novartis partnerships, and first clinical trials of AI-designed drugs described as imminent as of April 2026 (Source: wired.com); the company raised a $2.1B Series B led by Thrive Capital (May 13). Profluent and Eli Lilly announced a multi-program partnership for AI-designed recombinases — enzymes performing kilobase-scale DNA insertion — with up to $2.25B in milestones plus tiered royalties (April 28, 2026); recombinase engineering has historically been a manual, low-throughput process (Source: press.airstreet.com). University of Cambridge researchers reported the first human trial of a vaccine whose active antigen was entirely AI-designed, a "universal" sarbecovirus candidate targeting SARS, MERS, and Covid-19; the phase-one trial of nearly 40 people (2021–2023) recorded no serious side effects but only a "modest" immune effect, with a phase-two trial to follow (June 5, 2026). (Source: scmp.com) These complement the structure-prediction and drug-discovery threads on the development side, here with AI-designed immunogens reaching first-in-human testing. A Wall Street Journal Heard on the Street analysis published July 10, 2026 argued that AI has arrived in the drug lab but that drug discovery is not accelerating on Wall Street's investment timeline (Source: wsj.com).

Vendors and products

The two largest frontier labs launched dedicated healthcare products in early-to-mid 2026, positioned for different audiences. OpenAI offers "OpenAI for Healthcare" and "ChatGPT Health" (see OpenAI; (Source: openai.com), previewed Jan 7, 2026, with an official product page on May 15, 2026); ChatGPT Health is consumer-default (Free/Go/Plus/Pro tiers). Anthropic offers Claude for Healthcare ((Source: anthropic.com), Jan 12, 2026), positioned as enterprise (HIPAA-ready, for payers, providers, and healthtech), and also runs Claude in clinical-documentation workflows via Thomson Reuters and Bedrock. The split is described as an instance of OpenAI-consumer versus Anthropic-enterprise positioning: OpenAI prioritizes the consumer relationship while Anthropic prioritizes the enterprise workflow. The following table summarizes the two product architectures.

ArchitectureOpenAI ChatGPT Health ((Source: openai.com))Anthropic Claude for Healthcare ((Source: anthropic.com))
Launch dateJan 7, 2026 (preview) → May 15, 2026 (official product page)Jan 12, 2026
Default audienceConsumer (Free/Go/Plus/Pro)Enterprise (HIPAA-ready; payers, providers, healthtech)
Medical-records connectorb.wellCMS Coverage Database + ICD-10 + NPI Registry
Research connectorPubMed, ToolUniverse, bioRxiv/medRxiv
Wearables / personal appsApple Health, Function, MyFitnessPal, Weight Watchers (GLP-1)Apple Health, Android Health Connect, HealthEx, Function
Agent skills(limited; HealthBench evaluator-internal)FHIR development, prior-authorization review, clinical-trial protocol generation
Cloud footprintOpenAI infrastructureAWS Bedrock + Google Cloud Vertex + Azure (only frontier model on all 3)
Notable customer cited(none at launch)Advocate Health (167K-person system; full-scale Claude + PwC deployment) (Source: anthropic.com)
Physician collaboration260+ MDs / 60 countries / 600K+ feedback events; HealthBench evaluation"Detailed simulations of medical/scientific tasks"; MedCalc + MedAgentBench benchmarks
"Not for diagnosis or treatment" disclaimer
230M / week health Q's to ChatGPTself-disclosure anchor

On July 23, 2026, OpenAI opened ChatGPT Health to all US users over 18 on every plan, citing 300 million weekly health queries (up from 230 million in January) and adding medical-records access via Epic and Oracle Health; the launch came one day after a Florida pastor sued OpenAI alleging GPT-4o gave "extremely dangerous medical recommendations," including delaying treatment for a pulmonary embolism (Source: techcrunch.com; nytimes.com). On the federal side, HHS announced a series of AI healthcare challenges with OSTP on July 22, 2026 (Source: insideaipolicy.com).

Other vendors include Google DeepMind, whose AlphaFold (DeepMind) and AlphaGenome (Google DeepMind) support drug discovery and genomics, and Microsoft, whose DAX Copilot provides ambient clinical documentation. OpenEvidence serves as a widely adopted medical-search platform (see Adoption patterns).

In drug discovery, Anthropic launched an AI drug-discovery program built on its Claude Science workbench on June 30, 2026, joining other technology companies expanding into healthcare applications (Source: cnbc.com); Isomorphic Labs and Midjourney (whole-body ultrasound screening) are other AI-firm entries into the sector.

Product architecture detail

Claude for Healthcare's four connectors are the CMS Coverage Database (Local and National Coverage Determinations, positioned for prior-authorization checks and claims appeals), an ICD-10 connector covering diagnosis and procedure codes for coding, billing, and claims management, the National Provider Identifier Registry for provider verification, credentialing, and claims validation, and PubMed's 35 million-plus biomedical items, extended through HIPAA-compliant Claude for Enterprise. Two Agent Skills accompany them: FHIR development for cross-system interoperability, and a prior-authorization review template that cross-references coverage requirements, clinical guidelines, patient records, and appeal documents. Anthropic lists care coordination and patient message triage among intended use cases without publishing deployed-customer outcome data for them (Source: anthropic.com).

The "HIPAA-ready" claim rests on Claude for Enterprise as the deployment surface rather than on a legal certification; customers remain responsible for their own business associate agreements. A consumer-facing layer for Pro and Max users adds HealthEx and Function connectors in beta, with Apple Health and Android Health Connect integrations following, described as "private by design," opt-in, and excluded from training (Source: anthropic.com).

The Claude for Life Sciences expansion, first announced in October 2025, added Medidata (historical enrollment and site performance for Study Feasibility customers), ClinicalTrials.gov, ToolUniverse (a library of more than 600 vetted scientific tools), bioRxiv and medRxiv, Open Targets for therapeutic drug-target prioritization, ChEMBL, and Owkin Pathology Explorer for tissue-image analysis, joining existing Benchling, 10x Genomics, BioRender, Synapse.org, and Wiley Scholar Gateway connectors. New Life Sciences Agent Skills cover scientific problem selection, instrument-data-to-Allotrope conversion, scVI-tools with Nextflow deployment for bioinformatics, and clinical-trial protocol draft generation accounting for FDA and NIH requirements, endpoint recommendations, competitive landscape, and regulatory pathways. Listed implementation partners are Accenture, Blank Metal, Caylent, Deloitte, Deepsense.ai, Firemind, KPMG, Provectus, PwC, OWT, Quantium, Slalom, Tribe AI, and Turing (Source: anthropic.com).

Anthropic's benchmark claims for the offering rest on MedCalc, a medical-calculation-accuracy benchmark with Python execution, and MedAgentBench, a Stanford medical-agent task set, stating that Opus 4.5 "represents a major forward step" on medical benchmarks. Both are cited without independent verification, as is the accompanying claim that Opus 4.5 reduces "factual hallucinations" in honesty evaluations, referenced to the system card; no independent audit of clinical-context hallucination rates is provided (Source: anthropic.com).

On the OpenAI side, ChatGPT Health is compartmentalized — separate memories, conversations, and files, with information not flowing from Health back into non-Health chats — and OpenAI states that Health conversations are not used to train its foundation models. Its HealthBench evaluation uses physician-written rubrics scoring safety, clarity, escalation appropriateness, and respect for individual context rather than exam-style accuracy, built from more than 600,000 feedback events from 260-plus physicians across 60 countries and 30 specialties over two years. The launch page cites no peer-reviewed clinical-outcome studies, a gap Topol/Marcus: LLMs and Patient Outcomes — three-document cluster addresses in finding "very little evidence for LLMs benefiting patients or doctors for health outcomes" outside administrative work (Source: openai.com).

Risks and oversight

Deskilling. Lancet Endoscopist Deskilling Study (2025) is described as the first peer-reviewed clinical-endpoint evidence of AI-induced deskilling: the adenoma detection rate on colonoscopies performed without AI fell from 28.4% before AI was introduced to 22.4% afterward, a six-percentage-point decline, even though AI-assisted colonoscopies improved detection (N=1,443). The study assessed colonoscopy quality in the three months before and after AI implementation, indicating that exposure to the tool eroded unassisted clinical skill (Source: luizasnewsletter.com). See AI Deskilling. The AMA survey finding that 88% of physicians are concerned about skill loss (AMA Physician AI Sentiment Report (2026)) reflects the same concern at the level of professional attitudes.

Patient-outcomes evidence gap. Eric Topol published a May 3, 2026 review concluding there is "very little evidence for LLMs benefiting patients or doctors for health outcomes" outside administrative work (clinical-note drafting, prior-authorization paperwork, insurance-coding workflows) (Topol/Marcus: LLMs and Patient Outcomes — three-document cluster). Gary Marcus echoed the point the same day, citing a Nature Medicine editorial. (Source: garymarcus.substack.com) Topol and Marcus argue that AI value is concentrated in administrative workflows — prior authorization, claims appeals, billing, ambient scribing — where the empirical evidence is best-established, while clinical decision-making remains in support-not-replace territory; this distinction recurs as the reference point for healthcare-AI return-on-investment claims, and pairs with the Lancet endoscopist-deskilling study as counter-evidence to claims that AI is transforming clinical care. See Eric Topol, Gary Marcus.

Hallucination in clinical settings. Reporting under the headline "AI failed to detect critical health conditions: study" documents the failure mode of AI missing diagnoses. Hallucination risk in medical contexts is treated as qualitatively different from other domains because wrong medical information can lead to patient harm. See Sycophancy and Hallucination.

Mental-health applications. MIT Technology Review reported (Sept 2025) on therapists secretly using ChatGPT during sessions, and Nextgov reported (Oct 2025) on VA AI suicide-prevention tools described as not meant to replace clinical interventions. The mpathic benchmark (May 12), described as the first clinician-built mental-health AI benchmark, tested 6 leading models on suicide risk, eating disorders, and misinformation, finding that the most common harmful behavior was "reinforcement," in which models validate user beliefs without scrutiny. Online therapy platform Headway began requiring clients and providers to undergo biometric facial scanning, via third-party vendor Persona (a Founders Fund portfolio company), for identity verification; per a Headway email to clients on April 3, 2026 (surfaced by 404 Media on May 28, 2026), there is no opt-out other than leaving the platform, the requirement rolls out first to patients of prescribers, sessions are "auto-cancelled if verification is incomplete," and Headway is coaching providers on how to convince wary clients to comply. The report describes Headway as the first major US telehealth platform to make biometric verification a non-negotiable condition of continued care, with relevance to the AI mental health vendor space, the AI surveillance thread, and the identity-verification economy described in Persona (stub). (Source: 404media.co) See also Companion Chatbot Harms — Cross-Cutting Analysis and AI Psychosis. A Wall Street Journal investigation published July 11, 2026 documented the reinforcement failure mode in eating-disorder care specifically: AI chatbots dispensing reasonable-sounding nutrition and fitness guidance were undermining eating-disorder therapy, with clinicians reporting chatbot advice that reinforced patients' disordered behaviors (Source: wsj.com).

Regulation. The FDA SaMD framework is the primary US regulator for AI in clinical practice. The FDA AI/ML Action Plan (ongoing since 2021) sets out the SaMD regulatory framework; the agency completed an AI-Assisted Scientific Review Pilot (May 8, 2025) and deployed an agency-wide AI tool for internal performance optimization (June 2, 2025). California AB 3030 (Healthcare AI Disclosure) requires a generative-AI disclaimer for healthcare AI in California. AMA Physician AI Sentiment Report (2026) reflects professional-body positioning. Under the EU AI Act, healthcare AI is classified as high-risk under Annex III. Internationally, Brazil's Federal Council of Medicine (CFM) published Resolution No. 2,454/2026 in the Diário Oficial da União on February 27, 2026, effective 180 days later, establishing a sector-specific standard of care for AI in medicine: a life-cycle approach covering research, development, governance, auditing, monitoring, and training, and requiring human oversight, transparency, and data-protection safeguards when AI systems — including large language models and generative tools — process sensitive health data (Resolução CFM nº 2.454/2026 — regulation of AI use in medicine (Brazil, February 2026); Source: iapp.org).

The resolution's operative rules place the physician as final decision-maker in all diagnostic, therapeutic, and prognostic decisions, and separately prohibit delegating to AI "the communication of diagnoses, prognoses or therapeutic decisions" — so a system may inform a decision but the disclosure to the patient may not be automated. Physicians may refuse to use tools that are not scientifically validated, lack regulatory certification, or conflict with professional ethics, and may reject a system's recommendation "without suffering penalization." Liability protection is conditional rather than automatic: physicians are shielded from responsibility for failures attributable exclusively to AI systems "provided that diligent, critical and ethical use of the tool is demonstrated," with corresponding duties to exercise critical judgment, track the systems' limitations, and record the use of AI in the medical chart. Systems are classified across four risk levels — low, medium, high, and unacceptable — by reference to impact on fundamental rights, model complexity, degree of autonomy, and data sensitivity. Institutions running their own systems must establish internal governance and, where applicable, a Commission on AI and Telemedicine under medical coordination; enforcement runs through the Regional Councils of Medicine rather than an administrative regulator (Resolução CFM nº 2.454/2026 — regulation of AI use in medicine (Brazil, February 2026)).

The Centers for Medicare and Medicaid Services (CMS) launched ACCESS, an outcome-based payment model, with 150 health-AI companies in the first pilot, requiring vendors to demonstrate measurable patient impact rather than running demonstrations for hospitals (May 13). The 10-year program goes live July 5, 2026 and creates the first Medicare mechanism to pay for AI agents that monitor patients between visits. (Source: techcrunch.com)

The appropriate degree of clinical autonomy for AI surfaced as a point of disagreement between regulators and the profession. On June 22, 2026, CMS Administrator Mehmet Oz urged the American Medical Association to be open to autonomous and semiautonomous AI in health care, while AMA CEO John Whyte countered that AI should support rather than replace clinician decision-making and needs new guardrails (Source: insideaipolicy.com). The exchange tracks the support-not-replace framing reflected in the AMA's own survey data (AMA Physician AI Sentiment Report (2026)).

The FDA also began weighing AI in drug review. At a public workshop reported June 22, 2026, the agency and industry stakeholders considered using AI to make generic-drug development and review more efficient, while flagging unresolved questions about validating the tools, guarding against errors, and defining human accountability (Source: insideaipolicy.com). See FDA — Food and Drug Administration (AI Deployer).

The FDA's draft agreement on Medical Device User Fee Amendments (MDUFA) reauthorization, cleared by the Office of Management and Budget on July 8, 2026, includes most digital-health priorities stakeholders sought, including health-technology sandboxes and a streamlined pre-submission process (Source: insideaipolicy.com).

In Congress, Rep. Michael McCaul introduced the Accelerating Innovation for Kids with Cancer Act on July 9, 2026, which would create an HHS AI-innovation coordinator for pediatric cancer research (Source: nextgov.com).

Cybersecurity. Healthcare advocacy groups, including the Healthcare Leadership Council, called on July 8, 2026 for federal cybersecurity funding and guidance after the Commerce Department's late-June lifting of export controls on Anthropic's Mythos 5 and Fable 5, models that raised concerns over their ability to identify and potentially exploit cyber vulnerabilities (Sources: insideaipolicy.com; politico.com). See Export Controls (AI), Claude Mythos 5.

Brazil's resolution is tracked as a legislative instrument at Resolução CFM nº 2.454/2026 (Brazil — AI in medical practice).

OpenAI's evaluation framework is tracked at HealthBench, including the distinction between what its physician rubrics measure — safety, clarity, escalation appropriateness, and respect for individual context — and the clinical-outcome evidence they do not supply.

Relationships