AI is used in healthcare in twelve recognisable ways today, spread across five areas: clinical decisions, documentation, patient-facing services, hospital operations, and research. Some of those uses are regulated as medical devices by the FDA. Some are governed as predictive decision support inside certified electronic health record (EHR) systems under the federal HTI-1 rule. Some sit outside both and are bound only by HIPAA. That distinction decides more about a project than the technology does.
I lead AI transformation work at AppVerticals, and the question I get from healthcare founders and CTOs is rarely “what can AI do”. It is “which of these is real, and what would building one commit us to”. So this is a map rather than a highlight reel: what each use case actually does, what the published evidence shows and does not show, and what has to exist before any of it can reach a patient.
Key Takeaways
- Physician use of AI reached 81% in the American Medical Association’s 2026 survey, up from 38% in 2023, with an average of 2.3 use cases per physician.
- Every healthcare AI use case falls into one of three regulatory lanes: an FDA-regulated medical device, a predictive decision support intervention inside certified health IT, or unregulated operational software. The lane decides your obligations before a line of code is written.
- Health IT certified to the HTI-1 decision support criterion must publish 31 source attributes for predictive decision support and 13 for evidence-based decision support, covering how a model was trained, how it performs, and how it is maintained.
- Ambient documentation carries the strongest deployment evidence of any current use case: a 2% absolute reduction in burnout prevalence at Mass General Brigham at 84 days, published in JAMA Network Open.
- That same evidence base has real limits. A parallel Sutter Health study found note time fell from 6.2 to 5.3 minutes per appointment, while its burnout change did not reach statistical significance.
- Nothing ships without five prerequisites: usable structured data, an agreement covering the specific model endpoint, an integration surface, a human review checkpoint, and post-deployment monitoring.
What Does AI Actually Do in Healthcare? Analytic, Generative and Agentic AI Explained
Three kinds of AI are in play, and they behave differently enough that treating them as one thing causes most of the confusion.
Analytic AI learns patterns from historical data and produces a number: a probability, a classification, a measurement. It has been in hospitals for decades, reading electrocardiograms and flagging screening results. It is narrow, testable, and comparatively easy to validate because you can check its output against a known answer.
Generative AI produces new text, images or audio from a prompt. Large language models, systems trained on enormous volumes of text to predict what comes next, belong here. They are what changed the conversation in 2023, because for the first time software could draft a clinical note that read like a person wrote it.
Agentic AI takes the generative layer and gives it the ability to plan a sequence of steps and act on systems rather than only produce output. In healthcare it is the least mature of the three, mostly appearing in back-office workflows where an error costs money rather than harm. If you want the general mechanics, we covered agentic AI separately.
The practical distinction: analytic AI predicts, generative AI drafts, agentic AI does. The further right you move, the more the review checkpoint matters, because the software is taking more of the action away from the person.
AI in Clinical Decisions and Diagnostics
This is the category everyone pictures, and it splits into three distinct jobs.
- Diagnostic and imaging support. Algorithms read radiology, pathology and ophthalmology images and flag findings for the clinician reviewing them. Screening programmes such as mammography are where these tools have the most deployment history, because the review volume is high and the target is well defined. We go into how these systems are built and validated in our guide to how AI supports medical diagnosis.
- Acute triage and prioritisation. Here the model does not diagnose; it reorders a queue. When a scan is taken, the algorithm looks for a small set of time-critical findings, a large vessel occlusion, a pulmonary embolism, and pushes that study to the top of the worklist. The clinical value is minutes saved before a human has even opened the study.
- Risk prediction and deterioration alerts. These models watch data already flowing through the EHR, vitals, labs, medication orders, and produce a score for a future event such as sepsis or readmission. They are the most common form of AI running quietly inside hospital systems, and also the most controversial, because a risk score that performs well in one population can perform badly in another.
The pattern across all three: the model narrows attention. It does not close a decision.
How AI Is Used in Clinical Documentation and Ambient Scribing
- Ambient clinical documentation. An application listens to the consultation with the patient’s permission and drafts the clinical note, which the clinician then reviews, edits and signs into the record. This is the use case with the strongest published evidence behind it, and the one most likely to be live at a health system near you.
- Chart summarisation and inbox drafting. The same underlying models condense a long patient history into a pre-visit summary, or draft a first reply to a patient message for the clinician to approve. The work is lower-stakes than the note itself, which is why it spread quickly.
The reason this category moved faster than anything clinical is that the problem it attacks is enormous and non-clinical. Documentation load is one of the structural drivers of clinician burnout, and it accumulates in the evening after the clinic closes.
What the strongest study found: A JAMA Network Open study led by Mass General Brigham researchers, surveying more than 1,400 physicians and advanced practice providers across Mass General Brigham and Emory Healthcare, found a 21.2% absolute reduction in burnout prevalence at Mass General Brigham at 84 days, and a 30.7% absolute increase in documentation-related wellbeing at Emory at 60 days. Rebecca Mishuris, MD, MPH, MS, Chief Medical Information Officer at Mass General Brigham and a co-senior author, reported that physicians described getting “their nights and weekends back.”
Hold that number loosely for now. I come back to what it does and does not establish further down.
Patient-Facing AI: Triage, Scheduling, Monitoring and Adherence
- Symptom triage and care navigation. Conversational tools take a description of symptoms and route the person toward the right level of care. Used well, they reduce unnecessary visits and shorten the path for people who genuinely need to be seen. Used badly, they become a liability surface, which is why the good implementations are conservative and escalate readily.
- Conversational scheduling and intake. This is how conversational AI is used in healthcare most often in practice: booking, rescheduling, reminders, insurance capture and pre-visit forms. It touches patient data without touching clinical judgement, which makes it one of the cleanest starting points for an organisation that has never deployed a model.
- Remote patient monitoring and adherence. Data from connected devices and wearables streams in continuously, and a model decides which readings deserve a human’s attention. Without that filter, remote monitoring generates more alerts than any care team can process. The AI is the thing that makes the programme survivable. The same architecture underpins much of what we build into telemedicine platforms.
How AI Is Used in Healthcare Operations and Revenue Cycle
Operational AI gets a fraction of the press and a large share of the deployments, because the failure modes cost money rather than harm.
- Medical coding and claims. Models read the clinical documentation and propose the billing codes, with humans handling exceptions instead of first-pass work. Coding is a natural fit because the input is text, the output is a constrained code set, and the correct answer is auditable after the fact.
- Prior authorisation and utilisation review. Systems assemble the documentation a payer requires and draft the submission. This is one of the fastest-growing uses, and one where the governance questions are sharpest, because automation sits between a patient and an approval.
- Capacity, staffing and scheduling optimisation. Forecasting models predict admissions, theatre utilisation and staffing need. This is old-fashioned analytic AI doing unglamorous work, and it is frequently the highest-return project on the list.
Operational systems in healthcare carry patient data even when they never touch a clinical decision, and the compliance obligations follow the data rather than the clinical stakes. A platform we built for Collaborative Patient Care Group let offshore representatives operate unattended kiosks inside medical supply stores and healthcare facilities, no clinical function anywhere in it, and every control that applies to patient documents still applied. We covered that build in detail in our HIPAA build guide, where compliance splits across three owners.
Ben Shahshahani, PhD, Chief AI Officer at Cleveland Clinic, describes the shift in blunt terms: “AI is no longer an experiment.” Operational deployment is most of what that sentence is describing. If you are evaluating adjacent systems, our guide to healthcare CRM systems covers the same territory from the platform side.
AI in Medical Research and Drug Discovery
- Trial matching, biomarker discovery and literature synthesis. Models scan patient records against trial eligibility criteria, sift genomic and proteomic data for candidate biomarkers, and compress research literature into something a working scientist can actually read.
Research applications sit outside clinical care delivery, which changes the risk picture entirely. A model that proposes a wrong biomarker candidate wastes laboratory time. A model that proposes a wrong diagnosis harms a person. That asymmetry is why research AI has been allowed to move faster.
How Is AI Used in Healthcare Today? What Adoption Actually Looks Like
Adoption is now high enough that “are hospitals using AI” is the wrong question. The American Medical Association’s 2026 Physician Survey on Augmented Intelligence, based on responses from 1,692 physicians surveyed in early 2026, found 81% of physicians using AI in a professional capacity, up from 66% in 2024 and 38% in 2023. The average physician using AI reported 2.3 use cases.
AMA Chief Executive Officer John Whyte, MD, MPH, framed the finding directly: “AI has quickly become part of everyday medical practice.”
Two caveats belong next to that number, and they matter more than the number does.
First, the survey measures individual physicians, not organisations. A hospital counts toward the adoption picture the moment one clinician uses one tool. End-to-end institutional adoption is considerably lower than 81% implies.
Second, the same survey asked what physicians require before adopting a tool, and the answers read like a product requirements document: validated safety and efficacy from a trusted source, strong data privacy protections, a feedback channel when something goes wrong, coverage under standard malpractice insurance, and EHR integration that works. Peer recommendation ranked well below all of them.
Read that list again if you are building. Physicians are not asking for better model performance. They are asking for evidence, contracts, escalation paths and integration. Four of those five are engineering and legal problems, not machine learning problems.
The Risks: Bias, Hallucination and Who Stays Accountable
Three failure modes recur, and each has a specific engineering response.
Bias from unrepresentative training data. A model trained mostly on one population performs worse on others, and in medicine that gap is a safety problem rather than a fairness abstraction. The response is local validation: test the model on patients resembling yours before deployment, and keep testing after.
Hallucination. Generative models produce fluent output whether or not the underlying content is correct, and fluency is persuasive. The response is grounding — constraining the model to retrieve from your own verified source material — plus a human checkpoint that is enforced by the interface rather than by policy.
Accountability drift. When a model runs quietly inside a workflow for long enough, the review step becomes a formality. The response is to build friction that survives familiarity: require an affirmative edit or sign-off, log it, and monitor how often the human actually changes anything. A review rate that trends toward zero is a warning, not a success metric.
None of these are solved. They are managed, and the managing is a permanent operating cost rather than a launch task.
The Three Regulatory Lanes AI in Healthcare Falls Into
Here is the part that gets decided too late on most projects. Before you evaluate a model, work out which of three lanes your use case sits in, because the lane sets the obligations, the timeline and most of the budget.
| Lane | What it covers | Regulator / rule | What the builder owes | Worked example |
|---|---|---|---|---|
| 1. FDA-regulated medical device | Software that informs a diagnosis or treatment decision for an individual patient | FDA, via the 510(k), De Novo or PMA pathways | Marketing authorisation, clinical evidence, a change-control plan for model updates, post-market surveillance | An imaging algorithm that flags a suspected large vessel occlusion |
| 2. Predictive DSI inside certified health IT | A model producing a prediction, classification, recommendation, evaluation or analysis, supplied through a certified EHR | ASTP/ONC, via the HTI-1 decision support intervention criterion | 31 source attributes describing development, performance and maintenance; intervention risk management covering risk analysis, mitigation and governance | A sepsis risk score surfaced inside the EHR |
| 3. Unregulated operational software | Everything that touches patient data without informing a clinical decision | HIPAA Security Rule only | A Business Associate Agreement covering the specific service, access controls, audit logging of the AI path | A claims-coding assistant or scheduling agent |
Lane 2 is the one most teams have never heard of. The HTI-1 final rule created a decision support intervention criterion in the ONC Health IT Certification Program — the first substantial revision to certified decision support requirements since 2012, and the federal government’s first attempt to regulate healthcare AI outside FDA-regulated devices.
Under it, certified health IT supplying a predictive decision support intervention must make available 31 source attributes across nine categories, and 13 for evidence-based decision support. Those attributes cover what data the model was trained on, how its performance was measured, how it is maintained, and on what schedule it is revalidated for fairness. They exist so an organisation can judge whether a tool is fair, appropriate, valid, effective and safe — the criteria the rule abbreviates as FAVES. They must be written in plain language, not technical documentation.
Lane 1 has a public verification tool most buyers never use. The FDA maintains an AI-Enabled Medical Device List of devices authorised for marketing in the United States. If a vendor makes a clinical claim, the product should appear there. The FDA is explicit that the list is not comprehensive, so absence is a question to ask rather than a verdict — but presence is verifiable in thirty seconds.
My take on why this matters more than the model choice: Two teams can build the same sepsis alert. One surfaces it through a certified EHR module and inherits the source attribute obligations, the risk management documentation and the revalidation schedule. The other runs it as an internal analytics dashboard and inherits none of them. Same maths, different companies afterwards. Work out the lane in week one, not in the security review.
Put a rough number on the idea.
Scope a first estimate against feature set and platform before the conversation gets formal.
Estimate the cost of your scopeWhat the Evidence on AI in Healthcare Shows, and Where It Does Not
The documentation studies above are the best evidence in the field, and it is worth being precise about what they establish.
The Mass General Brigham and Emory study surveyed pilot users who had opted in. At Emory, all 557 pilot users were surveyed and the response rate was 11%. The authors state plainly that their findings likely represent the experience of more enthusiastic users, and that broader study is needed before the results can be generalised to a full clinician population.
A separate JAMA Network Open study at Sutter Health found mean note time per appointment fell from 6.2 minutes to 5.3 minutes, a statistically significant result. In the same study, burnout fell from 42.1% to 35.1%, and that change did not reach statistical significance.
Both of those results are real, and they say different things. The time saving is measurable and replicated. The burnout effect is promising and, in at least one large deployment, not yet demonstrated to a conventional statistical standard. A business case built on the first is sturdier than one built on the second.
Two more honest limits are worth naming. Device authorisation is not the same as clinical evidence: independent reviews of FDA AI device summaries have found that randomised trial data, patient-outcome reporting and demographic transparency each appear in only a minority of authorisation records. And vendor-reported accuracy is almost always development-set performance, which tells you the model learned its training data, not that it will work on your patients.
None of that argues against deploying. It argues for asking the second question after the demo: show me the validation on a population like mine.
What Has to Be True Before an AI Use Case Can Ship in Healthcare
Five things, in build order. Projects that skip any of them do not fail at the model; they fail three months later at integration or at the security review.
- Data that a model can actually use. The features above assume clean, structured, retrievable clinical data. Where the organization’s data already conforms to FHIR R4 and the United States Core Data for Interoperability (USCDI), the standard formats for exchanging health records, the integration is comparatively cheap. Where it does not, most of the budget goes into extraction and normalisation before any model work starts. This is the single most common reason a healthcare AI timeline slips.
- A contract that covers the specific endpoint. Any external model that processes Protected Health Information requires a Business Associate Agreement naming that service and those endpoints. Consumer and enterprise tiers of the same product are different contractual instruments. Assuming coverage because you have an agreement with the parent company is a reliable way to fail an audit.
- An integration surface that exists. The model has to receive data from, and return output to, the systems clinicians already work in. EHR vendors control that surface, and their permission model, API limits and certification requirements shape what is buildable. This is ordinary system integration work, and it is usually the largest line item.
- A human checkpoint the interface enforces. Decide before build where a person reviews the output, and make that step structural rather than advisory. Log every review. If nobody can point to the checkpoint in the UI, it does not exist.
- Monitoring that runs after launch. Model performance degrades as clinical practice, coding conventions and patient populations shift, drift. Ship with a revalidation schedule, a defined performance floor and a rollback path. The HTI-1 source attributes explicitly ask developers to describe their maintenance and revalidation schedule, which is a reasonable standard to hold yourself to even in lane 3.
Where to Start: Sequencing Healthcare AI Use Cases by Risk and Time to Value
Score each candidate on three axes, low score meaning easier. Add them. The lowest totals are where a first project belongs.
| Use case | Regulatory burden | Integration depth | Time to value | Total |
|---|---|---|---|---|
| Conversational scheduling and intake | 1 | 2 | 1 | 4 |
| Capacity and staffing optimisation | 1 | 2 | 2 | 5 |
| Chart summarisation and inbox drafting | 2 | 3 | 1 | 6 |
| Medical coding and claims | 1 | 3 | 2 | 6 |
| Ambient clinical documentation | 2 | 4 | 1 | 7 |
| Trial matching and research synthesis | 1 | 3 | 4 | 8 |
| Remote patient monitoring and adherence | 3 | 4 | 3 | 10 |
| Prior authorisation and utilisation review | 3 | 4 | 3 | 10 |
| Symptom triage and care navigation | 4 | 3 | 3 | 10 |
| Risk prediction and deterioration alerts | 4 | 4 | 4 | 12 |
| Acute triage and prioritisation | 5 | 4 | 4 | 13 |
| Diagnostic and imaging support | 5 | 4 | 5 | 14 |
The scores are mine and they are a starting point, not a verdict, re-score them for your own context. An organization that already runs a certified EHR with a working integration layer should drop integration depth by a point or two across the board, which reorders the middle of the table substantially.
The shape holds regardless of how you re-score it. Work in lane 3 clusters at the top. Lane 1 clusters at the bottom. A first project in lane 3 teaches an organization how to govern a model, validation, review, logging, monitoring, while the consequences of getting it wrong are a wasted week rather than a harmed patient. That capability transfers. The model does not.
If you are working out what a build in any of these lanes involves, our healthcare app development work covers the engineering side in detail.
Visual 3 — case-study sidebar card. BLOCKED. Do not build this card until a verified healthcare-AI project with PM-confirmed outcomes is supplied. Do not populate it with CPCG, which anchors the HIPAA guide.
Conclusion
The twelve use cases in this article are all real, and they are not equally available to you. What separates them is the lane each one sits in and the work that has to happen before the model is even the interesting part. A health system with clean FHIR data and a signed agreement covering its model endpoints can move on documentation support in a quarter. The same organisation attempting an FDA-cleared diagnostic starts a multi-year programme with clinical evidence at its centre.
So the useful next step is narrow. Take the one use case your organization keeps returning to, put it in a lane, and find out what that lane asks of you. Most of what follows, budget, timeline, who needs to be in the room, falls out of that answer.
Before you brief anyone, work out what you owe
The compliance obligations on a healthcare build split across your cloud platform, your code and your own organization. This guide maps who owns which control.
Read the HIPAA build guide
ChatGPT