Artificial intelligence (AI) is moving from documenting care to participating in it. Regulation must set the floor, but informed physicians will determine whether the transition is safe.
Healthcare AI is no longer confined to predicting a risk score or drafting a note.
The next generation of systems can retrieve a chart, summarize years of records, identify missing information, propose a plan, call other software tools, and initiate work across clinical, administrative, and financial systems. Oracle, for example, publicly describes a Clinical AI Agent product direction spanning chart review, documentation, coding, scheduling, draft orders, follow-up, and prior-authorization workflows; Oracle cautions that some described features are planned, may change, and are not all currently offered.[1]
The potential is broader than documentation.
Used well, these systems can reduce clerical work, expose information buried across fragmented records, help clinicians prepare for complex visits, match patients to trials, identify omissions, and give physicians more time for the part of medicine that cannot be automated: judgment, explanation, and human presence.
Used poorly, the same systems can amplify stale chart errors, obscure where patient data traveled, produce a persuasive answer without a trustworthy basis, and move a mistake through the healthcare system faster than a person could.
That is the central regulatory problem. We are not simply introducing a new medical device. We are introducing a class of systems whose behavior depends on the model, the prompt, the data retrieved, the tools connected, the workflow, the user, and the environment in which the output is acted upon.
AI capability and deployment are accelerating. New model releases, prompts, retrieval systems, and connected tools can alter behavior on software timelines. Policy and regulation cannot keep pace with that iteration cycle, and there is no realistic reason to expect them to. They require evidence gathering, public process, legal authority, and implementation. That mismatch is structural, not a temporary failure of regulators.
The answer is not weaker regulation. Regulation should establish an enforceable floor. Physician education and institutional governance must operate in the gap, at the speed of deployment.
The Food and Drug Administration has started asking the right questions
On August 18, 2026, the U.S. Food and Drug Administration (FDA) released Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. Comments are due October 19, 2026, under docket FDA-2026-N-7874.[2]
The legal status matters. This is not a proposed rule or draft guidance, and it creates no new regulatory requirement. FDA is seeking early feedback on possible approaches to generative-AI-enabled medical devices without deciding whether those approaches fit within its existing legal authority or would require new authority.[2][3]
The paper’s most useful idea is that risk should not be defined by the label “AI.” It should be defined by what the final product actually does.
A system that retrieves source-linked medical literature is not the same as a system that recommends a patient-specific treatment. A tool that drafts an order for review is not the same as one that places the order. A note generator is not the same as an autonomous agent controlling another device.
The paper explores a possible risk framework with two broad dimensions: how independently the device function acts and the consequences of relying on an incorrect output. Under that possible framework, evidence and controls could increase as a function moves from supplying information to directing action, taking action under supervision, or operating autonomously, and as the consequence of error rises. A flawed scheduling suggestion and a flawed dose calculation do not belong in the same category.[3]
That is a sensible start.
The FDA also recognizes that traditional validation is not enough. A conventional prediction model can often be tested against a defined input set and measured with sensitivity, specificity, calibration, or another bounded metric. A generative system can accept open-ended language, produce different responses to equivalent questions, change behavior across a long conversation, retrieve external information, and act through connected tools.
Exhaustively testing every possible interaction is not practical.
The FDA’s Center for Devices and Radiological Health (CDRH) is therefore considering a competency-based approach that would combine non-clinical device benchmarking with clinical confirmation in real or clinically representative use. Possible methods discussed in the paper include adversarial testing, retrospective evaluation, specialist adjudication, standardized-patient interactions, silent shadow deployment in which outputs cannot affect care, and prospective studies when warranted by risk.[3]
The paper also treats post-deployment monitoring as a central question. Prompts change. Retrieval systems change. Guardrails change. The upstream general-purpose model may change. New tools are connected. A product that was safe at authorization can become meaningfully different without changing its name or interface.
Under the existing statutory and authorization framework, a manufacturer may implement certain planned modifications covered by an FDA-authorized predetermined change control plan (PCCP) without obtaining separate authorization for each modification; FDA’s final PCCP guidance provides nonbinding recommendations for such plans.[4] Generative AI complicates that framework because some important changes may originate upstream or emerge from interactions among several components.
The questions physicians should ask now
Physicians do not need to become software engineers. They do need enough technical fluency to understand how an AI system changes the clinical workflow and which questions affect patient safety:
What is the intended use? What is the system designed to do, and what is explicitly outside scope?
What information does it use? Which chart fields, notes, images, audio, external sources, and inferred data enter the workflow?
Where does the information go? Which vendors, downstream service providers, regions, logs, memories, retrieval stores, and backups receive a copy?
How was it validated here? Was the actual configuration tested in the local workflow, patient population, specialty, and clinical language?
Can the physician inspect the basis? Are claims linked to source records, or does the system return a polished conclusion without provenance?
What happens when the system is uncertain or wrong? Does it defer, stop safely, request missing information, or continue confidently?
What can it do without confirmation? Can it write, order, send, schedule, code, bill, or contact a patient?
How is change controlled? Are model, prompt, retrieval, tool, and interface versions recorded and revalidated?
How is performance monitored? Who reviews failures, what thresholds trigger rollback, and can the system be shut down quickly?
Who remains accountable? The answer cannot be “the AI.”
These questions are useful now because product deployment is already occurring while formal policy remains under development.
Why AI regulation is uniquely difficult
The FDA’s direction is promising, but healthcare AI does not fit neatly inside the regulatory structures built for static products.
1. The same model can occupy several regulatory categories
A foundation model is not regulated as a medical device merely because it is a foundation model. FDA’s discussion focuses on the final user-facing device function, including its intended use, configuration, and behavior, rather than on the foundation model standing alone.[3]
That distinction is essential. The same model could summarize a meeting, retrieve trial criteria, draft a patient message, recommend a diagnosis, calculate a treatment parameter, or execute an order. The risk does not come from the model name. It comes from the clinical function, the data, the degree of autonomy, and whether a qualified person can independently evaluate the basis.
FDA’s clinical decision-support guidance reflects the same principle. To qualify for the statutory non-device clinical decision support (CDS) exclusion, a software function must satisfy all four criteria in section 520(o)(1)(E) of the Federal Food, Drug, and Cosmetic Act (FD&C Act), including enabling a healthcare professional to independently review the basis of the recommendation rather than rely primarily on the software’s conclusion.[5]
A disclaimer does not create independent review. Neither does an “approve” button.
2. “Human in the loop” is not a sufficient safety claim
Human oversight is often treated as a binary property: either a physician clicks approve or the system acts autonomously.
That is too crude.
A physician cannot provide meaningful oversight if the source is hidden, the workflow is time-pressured, the reasoning cannot be reconstructed, or the interface makes verification slower than accepting the answer. In clinical practice, source review is not a theoretical safeguard. A fluent summary can produce more confidence than the underlying evidence deserves.
Clinical records make this especially important. The electronic health record (EHR) is not a clean representation of truth. It contains old problem-list entries, copied-forward notes, billing labels, conflicting histories, and conclusions written before later testing changed the diagnosis. An AI system that is “grounded in the chart” may still be grounded in the wrong part of the chart.
Real human oversight means review of the source, not merely review of the summary.
A published ambient-AI governance report illustrates how sensitive these systems can be to workflow changes: a documentation-template change was associated with International Classification of Diseases, Tenth Revision (ICD-10) documentation accuracy falling from 79% to 35% before redesign and user training restored performance.[6] The lesson is not that ambient AI is unsafe. It is that implementation, monitoring, and clinician education are part of the safety case.
3. Medical-device regulation does not cover the entire problem
FDA regulates device functions. It does not, by itself, answer every question about privacy, data residency, contracting, cybersecurity, professional accountability, or the quality of the underlying medical record.
The Health Insurance Portability and Accountability Act (HIPAA) requires covered entities and business associates to protect electronic protected health information with appropriate safeguards and contractual arrangements.[7] The Office of the National Coordinator for Health Information Technology (ONC) adds transparency expectations for certain predictive decision-support interventions in certified health information technology through its Health Data, Technology, and Interoperability final rule (HTI-1).[8]
Those are important layers. They are not a complete governance system for agentic clinical AI.
A business associate agreement does not tell a patient which copies of an encounter exist, how long audio or transcripts remain, whether a searchable memory persists, which downstream vendor handled the information, or whether optional tools fall inside the contracted environment.
One unusually transparent example comes from the Department of Veterans Affairs. Its privacy assessment for an Abridge ambient-scribe deployment describes audio and note data moving through Abridge’s Google Cloud environment, the clinician reviewing the draft before it enters the Veterans Health Information Systems and Technology Architecture (VistA), and the original vendor-held audio and written notes being deleted after a period of up to 30 days.[9]
A 2025 review found that residual re-identification risk in de-identified clinical free text is difficult to quantify, while concluding that properly de-identified text held in secure environments can carry very low risk.[10] De-identification is therefore a meaningful control, but not a substitute for understanding access, retention, and use.
That does not make cloud processing inherently unsafe. It demonstrates the level of specificity that should be normal.
“HIPAA-compliant” is not a data-flow diagram.
The transformation is larger than documentation
Ambient documentation is the entry point because it offers an obvious benefit: less time typing and more attention available for the patient. It is not the endpoint.
Healthcare AI is moving from systems that generate text to systems that coordinate work. An agent may retrieve results, compare prior notes, check a formulary, draft an order, prepare a prior authorization, send a message, or trigger another system. Every additional tool expands both capability and consequence.
A scribe can write the wrong sentence. An agent can use the wrong sentence to take the next step.
Properly designed AI could help clinicians navigate fragmented records, identify missing staging information, screen trial criteria, and detect when recommended follow-up never occurred. The governance question is whether that speed and scale can expand without making accountability less visible.
What regulation can address, and what it cannot
Formal regulation can establish requirements that the market may not produce consistently on its own.
It can tie evidence to the consequence of error, define expectations for change control and postmarket monitoring, and clarify responsibility when an upstream model or connected tool changes. It can also distinguish meaningful human review from a nominal approval step.
At the same time, every regulatory design involves tradeoffs.
If evidentiary requirements are not proportionate to risk, smaller clinical builders and independent practices may face costs that only large vendors can absorb. If requirements are too permissive, clinically important changes may reach patients without adequate revalidation. A framework focused only on foundation models may miss the final clinical function, while one focused only on premarket performance may miss workflow drift after deployment.
Architecture labels do not resolve the problem. Local deployment can improve control over versions and data, but it does not guarantee good security or validation. Cloud deployment can create additional data and vendor boundaries, but it can also operate under mature controls. The relevant evidence concerns the actual configuration, data flow, access, retention, validation, and monitoring.
Timing creates another constraint. Health systems are already adopting these tools while capabilities accelerate and the regulatory framework is still being developed. That gap cannot be closed by faster rulemaking alone. Formal oversight will therefore shape both new products and workflows that may already be installed.
These tensions explain why regulation is necessary but unlikely to function as the only near-term safety mechanism.
Why physician education matters during the transition
The same technical fluency belongs across the institution. Procurement teams encounter model and downstream-vendor boundaries. Privacy officers need retention schedules for each kind of data, not only general assurances. Clinical leaders need measures of human-AI team performance rather than adoption alone. Patients need a meaningful explanation of what the system does and how it affects their care.
Public vendor documentation already shows why product-specific diligence matters. Retention and improvement terms differ across ambient systems, and customer contracts may govern clinical data differently from public website policies.[9][11][12] The useful distinction is not trust versus fear. It is whether the architecture and terms for the exact deployed feature are visible enough to evaluate.
Four workstreams now developing in parallel
The current debate is not limited to one regulator or one technical standard. It spans four connected workstreams.
Questions for regulators
How should oversight follow the final clinical function rather than the general label “AI”?
How should evidence scale with autonomy, consequence, reversibility, traceability, and time available for review?
Which model, prompt, retrieval, tool, and version changes require documentation, revalidation, or a new submission?
How should accountability be divided between the final-product manufacturer and an upstream foundation-model provider?
Questions for health systems
Is there a feature-specific data-flow map for the exact product being deployed?
Has the actual configuration been evaluated in the local clinical workflow rather than only in a vendor demonstration?
Can clinicians inspect source-linked evidence before orders, messages, prescriptions, billing, or record changes occur?
Who owns ongoing sampling, incident review, rollback criteria, and shutdown authority?
Questions for developers
Are uncertainty, missing information, and conflicting source data visible to the user?
Can clinicians reconstruct the provenance of a clinically important claim?
Are prompts, retrieval data, tools, and model versions treated as controlled configuration?
Is evaluation measuring the human-AI team as well as the model in isolation?
Questions for physicians
Is the output being treated as a draft or as unquestioned chart truth?
Can the underlying source be checked before the output influences care?
Where does patient data go, how long does it remain, and can it be used for product improvement?
Are clinicians involved in validation, procurement, and post-deployment review?
The FDA is explicitly inviting comments from clinicians and researchers, not only manufacturers. The open docket gives physicians an opportunity to describe real workflow tradeoffs involving traceability, reversibility, review time, specialist use, shadow deployment, and the burden of ongoing monitoring. The comment period closes October 19, 2026.[2]
The emerging standard is informed accountability
The FDA’s discussion paper recognizes that existing evaluation methods may not be sufficient for generative and agentic systems and explores possible adaptations in evidence, monitoring, change control, and human oversight. Formal policy is still being designed while adoption is underway.
Education can be implemented faster than formal rulemaking, but it cannot create enforceable minimum standards. Regulation can create those standards, but it cannot substitute for local workflow knowledge or informed use.
Durable trust will therefore depend on how regulation, institutional governance, engineering controls, and physician education operate together.
When AI enters the exam room, the system should be visible enough to understand, bounded enough to supervise, monitored well enough to detect change, and connected to people who remain accountable for the care delivered with it.
Regulatory note: This article is an analysis of emerging policy, not legal advice. Device status and compliance obligations depend on the product’s intended use, design, deployment, contracts, jurisdiction, and actual clinical workflow.
Sources
U.S. Food and Drug Administration Generative AI-Enabled Medical Devices Discussion Paper
U.S. Food and Drug Administration Predetermined Change Control Plan Guidance
U.S. Food and Drug Administration Clinical Decision Support Software Guidance
U.S. Department of Health and Human Services: HIPAA Security Rule Summary
Office of the National Coordinator for Health Information Technology: HTI-1 Final Rule
Department of Veterans Affairs Abridge Ambient Scribe Privacy Impact Assessment
Ford et al.: Re-identification Risk from De-identified Clinical Free Text





