A 31-page discussion paper issued by the Digital Health Center of Excellence within the FDA's Center for Devices and Radiological Health, published August 19, 2026. Feedback is invited under docket FDA-2026-N-7874 on Regulations.gov by October 19, 2026; respondents may answer selectively and may provide partial responses (Source: fda.gov).
The paper's own front-matter disclaimer is unusually explicit and is reproduced rather than paraphrased. It "is intended for discussion purposes only and does not represent draft or final guidance," "is not intended to propose or implement policy changes regarding how CDRH intends to regulate generative AI-enabled devices," "is not intended to communicate CDRH's proposed (or final) regulatory expectations, including its expectations for supporting evidence in future marketing submissions," and "is not intended to address whether the approaches discussed below are within FDA's existing legal authorities or whether new legal authorities would be necessary." It is a request for feedback on four topics — risk assessment, premarket evaluation, postmarket monitoring, and other topics — carrying 26 numbered discussion questions, and is not a request for information on the regulation of generative AI generally.
Scope and definitions
The paper states plainly that "FDA does not regulate GenAI as such; it regulates medical devices, including GenAI-enabled devices," consistent with its treatment of software, hardware and AI generally. A GenAI-enabled device is a product containing one or more device software functions enabled by generative AI; where such a function is integrated into a broader device, the paper addresses the GenAI-enabled component only. Definitions are drawn from FDA's Digital Health and Artificial Intelligence Glossary and cover generative AI, foundation models (including emergent capabilities and fine-tuning), large language models, multimodal processing, and agentic AI systems, defined as "GenAI-enabled systems that autonomously plan and execute multi-step tasks, use external tools, or take actions across a sequence of steps." See Agentic AI.
The stated distinguishing characteristics of these devices are open-ended inputs, multiple subtasks, variable outputs to similar inputs, and evolution over time through changes to the underlying model, prompts, retrieval strategies, guardrails, orchestration logic or user interface. Because many are built on third-party general-purpose foundation models with varying transparency into training data, architecture and evaluation methods, the paper notes it can be "difficult to attribute specific behaviors and errors to the device or its underlying model." Named risks include confabulations that may appear authentic, uncertainty in the bounds of intended use, limited visibility into third-party models, and performance degradation across test and real-world environments.
The paper builds on CDRH's November 2024 Digital Health Advisory Committee meeting, which identified two categories of regulatory challenge: applying a risk-based approach to classification, and determining what constitutes valid scientific evidence across the total product life cycle.
The two-axis risk framework
The paper offers a two-axis framework as "a possible organizing heuristic," drawing structure from analogous matrices in other FDA guidance. The horizontal axis is device activity, increasing in degree and independence of action; the vertical axis is consequences, the severity of harm from relying on an incorrect output. Risk increases from the lower-left toward the upper-right.
Six considerations are set out under it:
- Action-directing functions. CDRH is considering treating informational functions that direct action as higher-risk than non-directive ones, while noting the distinction is a continuum rather than binary. The illustration runs from general information ("lisinopril dosages are sometimes increased when blood pressure remains above the treatment goal") through connecting general practice to circumstances, to endorsement, to specific instruction ("increase the lisinopril from 10 mg to 20 mg daily"). CDRH is considering that directiveness may depend on the substance and context of the output rather than on words such as "recommend," "should," or "consider," and that a patient-facing function may not become less directive by adding a "talk to your doctor" or "I am not a medical professional" statement.
- Action-taking functions. Assigning a clinical diagnosis, prescribing a medication or initiating a clinical order set is treated as typically higher-risk than an informational function, moderated by position on the consequences axis.
- Measurement and signal processing functions. These may be higher-risk despite yielding non-directive information, because users typically cannot independently evaluate the basis for the output.
- Patient-facing versus HCP-facing. The paper argues both sides: clinicians may be better equipped to recognize an incorrect output, which would shift patient-facing functions higher on the consequences axis, but there are "important benefits associated with the democratization of clinical information, with patient empowerment and engagement, and with avoiding medical parentalism that may underestimate patient capability."
- Generalist versus specialist HCP-facing. A function helping a generalist approach a specialist question may extend access to specialty knowledge, or present elevated risk where safe use depends on specialist contextualization.
- Multi-turn conversations and care escalation. CDRH is considering assessing conversational devices "across realistic conversational trajectories, not only at the level of individual functions," since a device may migrate from informational to action-directing over an exchange. For care escalation it is considering both directions of error: under-escalation risks delayed treatment, over-escalation risks anxiety, unnecessary tests and emergency utilization, resource burden, and erosion of trust that may reduce appropriate care-seeking over time.
Competency-based premarket evaluation
The premise is that evaluation methods built for software with bounded inputs and fixed outputs do not transfer, because "the range of possible inputs and outputs may be too large for such testing to be practical."
CDRH is considering an approach "inspired, at a high level, by how human clinicians are evaluated and credentialed," adapted for device regulation. The analogy is stated precisely: clinicians are not evaluated by exhaustive testing of every scenario, but through structured assessments of knowledge and reasoning (licensing and board examinations) plus supervised practice with progressively greater independence (clerkships, residency, fellowship), followed by ongoing assessment. Three published proposals are cited by name as prior art — Patel and Blumenthal in JAMA Health Forum (2026), Bergman, Wachter and Emanuel in JAMA (2026), and Freyer et al. in Nature Medicine (2025), the last raising accountability objections on the ground that human clinicians face professional, legal and reputational consequences and that analogous mechanisms are needed for flawed devices.
The approach has two parts, applied to "the final user-facing device, as configured and intended to be deployed for real-world use — and not the foundation model standing alone or other isolated subcomponent," and scalable via the two-axis framework.
Device benchmarking is conceived as scalable, high-throughput non-clinical evaluation. Its elements are grouped in four families, applied selectively by intended use and risk profile: Safety (S.1 safety-critical recognition and escalation; S.2 scope maintenance and boundary adherence; S.3 calibration, uncertainty communication and clinical deferral); Clinical Proficiency (E.1 clinical knowledge and task fidelity; E.2 information gathering and clinical analysis; E.3 quantitative and measurement analysis; E.4 communication quality and user comprehension); Generalizability (R.1 robustness, reliability and reproducibility; R.2 subgroup performance); and Agentic AI Capabilities (A.1, applicable only to agentic devices). CDRH notes that publicly available benchmarking assets may suffer from data contamination, saturation and lack of representativeness, and that sponsors might propose their own tailored tests. General principles include prespecified methods and acceptance criteria, scoring rubrics grounded in clinical guidelines or validated by qualified experts, and expert adjudicators "structurally independent from the device sponsor and, where the device incorporates a third-party model, from the developer of that model" — a requirement the paper states remains applicable "when the expert adjudicator is itself an LLM."
Clinical confirmation addresses what benchmarking cannot establish. It "might not require a prospective clinical study in every case." Five approaches are listed in approximate order of increasing rigor and patient exposure: retrospective evaluation on real patient inputs; shadow deployment, in which the device runs in a live clinical workflow on real patients but its outputs are not shown to clinicians or patients and do not affect care; standardized patient interactions using trained patient actors; clinician adjudication of real cases, blinded or unblinded; and prospective clinical study, in some instances a randomized controlled trial. Synthetic data modeled on real patient data is raised as a possible supplement where real inputs or subgroup samples are limited.
On standards of performance, the paper notes that for open-ended outputs "a single correct response often does not exist," and CDRH is considering comparison against a panel of qualified clinicians whose consensus reflects the standard of care, or a median clinician in practice, with the question being whether the device performs as well as or better than a reference clinician on the same task. It also raises comparing human-AI team performance against the device working autonomously.
Postmarket monitoring and change control
CDRH raises whether it may be appropriate "to accept greater premarket uncertainty regarding a GenAI-enabled device's benefit-risk profile through greater reliance on postmarket monitoring." Three approaches are named: periodic re-benchmarking against prespecified thresholds on a defined cadence and after triggering events including changes to the underlying model; periodic sample-based review by qualified independent clinician adjudicators; and performance degradation (drift) monitoring with specified analyses and thresholds. It asks whether machine-based supervisory agents could enable parts of this, and how the supervisory agent itself would be evaluated.
Change control distinguishes three kinds of post-deployment change: intentional discrete modifications by the sponsor; model-evolution changes occurring passively or incrementally because the device was designed to learn or adapt during use; and unplanned changes arising from updates to an underlying third-party foundation model. Oversight options range from documentation in the sponsor's quality management system to FDA authorization before implementation, with a Predetermined Change Control Plan named as one mechanism. The premarket competency assessment is proposed as the baseline against which a modified device could be "re-benchmarked."
Shared ecosystem responsibility is raised as an open question: manufacturers remain responsible for postmarket monitoring, but clinicians, patients, healthcare institutions, payers, professional societies, public-private consortia, standards-setting bodies and other federal and state authorities are named as potentially having roles, with question 21 asking how those roles can be structured "without diffusing manufacturer accountability."
Foundation Model Device Master Files and agentic systems
CDRH seeks feedback on voluntary Foundation Model Device Master Files, using the existing Device Master File program, under which foundation model developers and platform providers could submit structured model cards or system cards to FDA. Suggested content includes architecture, training data provenance and intended supported use cases; characterized behaviors, known limitations and failure modes relevant to healthcare; evaluation results on healthcare-relevant benchmarks including clinically meaningful subgroups; safety-relevant behavioral constraints and guardrails; update notification commitments; and audit log availability. Files would be held confidentially and referenced by sponsors with the holder's authorization; submission would not constitute authorization of the model for any intended use. Question 25 acknowledges the incentive problem directly, asking what would make the program useful "given that participation would be voluntary and that model developers may have limited incentive to disclose safety-relevant information."
On agentic AI, the paper notes deployment in care coordination, clinical documentation, patient outreach and clinical workflow support, "some or all of which may not be functions that are the focus of FDA's device regulatory oversight," but identifies a case that would be: an agentic system whose action sequences result in control of another medical device. Question 26 asks how "the elevated risk associated with autonomous multi-step action, tool use, and reduced opportunity for human review" should be reflected in acceptance criteria and oversight. This places a device regulator on record about agentic systems in a premarket-evaluation context.
Provenance
Fetched in full, 31 pages, from the sole media download link on the FDA's own landing page for the paper. Figures 1 and 2 are captioned in the capture but their graphical content is not reproduced.
Earlier coverage of the paper, carried into the wiki through a lede-only retrieval, described it as a "request for information" on "the regulation of generative artificial intelligence." Both terms are wider than the document: it is a discussion paper and request for feedback, on generative AI-enabled medical devices, and it disclaims being draft or final guidance. That coverage also glossed the premarket approach as technology "benchmarked for risk and then potentially shadowed side by side with human clinicians"; the paper's own terms are non-clinical device benchmarking followed by clinical confirmation, of which shadow deployment is one of five listed confirmation approaches.
Relationships
- depends-on: FDA — Food and Drug Administration (AI Deployer) — the issuing center's regulatory framework.
- deploys-in: Healthcare — AI Deployment — the deployment setting the paper regulates.
- related: Agentic AI — Section VII.B and benchmarking element A.1.
- related: AI Liability — the shared-ecosystem-responsibility and manufacturer-accountability questions.
- related: NIST AI Risk Management Framework 1.0, Algorithmic Accountability and Bias Audits.