Putting AI in a Hospital Without Losing Control of It

The first AI tool arrives in a hospital without a decision being made. A radiologist trials a free reading assistant. A resident starts drafting discharge summaries with a chatbot because it saves forty minutes. Somebody in finance uploads a spreadsheet of claim data to get a summary. None of it went through procurement, IT or the ethics committee, and by the time anyone notices, three different tools are handling patient data and nobody can say which.
This is how AI actually enters healthcare organisations. Governance that assumes a formal adoption decision is governing something that already happened.
What does AI governance in a hospital actually mean?
It means knowing which tools touch patient data, who is accountable for each decision they influence, and being able to show both.
It is not an ethics statement and it is not a ban. A ban produces exactly the shadow usage described above, with the added problem that nobody will tell you about it. The workable version is narrower and more practical:
- An inventory of AI tools in use, including the unofficial ones.
- A named clinical owner for each tool that touches care.
- A clear statement, per tool, of what it does and what it must not be relied on for.
- A record of what the tool said and what the human did about it.
- A review that actually happens, on a schedule, using the hospital's own data.
Where the accountability sits
With the clinician who acts, always — and a tool that blurs this is a tool to reject.
This is the load-bearing principle. A model that flags a finding is offering an input, exactly like a prior study or a lab result. The report carries a radiologist's name. The prescription carries a prescriber's. The discharge summary carries the signature of whoever signed it, and signing means having read it.
Two failure modes threaten this from opposite directions. Automation bias is the first: the human stops reading carefully because the machine has already looked. The second is the alibi, where a clinician defends a decision on the grounds that the system suggested it. Both are prevented by the same design choice — the human commits to their own assessment before seeing the model's, and the record shows they did.
The safe question is never "is the model right often enough". It is "when the model is wrong, who catches it, and how".
Which uses are low-risk and which are not
Sort tools by what happens when they are wrong, not by how impressive they are.
- Low risk: scheduling suggestions, queue forecasting, stock reorder proposals, coding assistance that a coder reviews. A wrong answer costs efficiency and is caught in the ordinary course of work.
- Medium risk: documentation drafting, triage ordering, summarisation of a record for a clinician who will also read the source. A wrong answer wastes time or misleads briefly, and the human is positioned to notice.
- High risk: anything producing a finding, a diagnosis, a dose, or a prioritisation that determines who is seen first. A wrong answer can reach a patient.
The governance effort should be concentrated on the third category, and the second category deserves more attention than it usually gets — a fluent, confidently wrong summary is more dangerous than an obviously broken one, because nothing about it invites checking.
The data question, which is now a legal one
Any tool that receives patient data is a processor acting on the hospital's behalf, and the hospital remains responsible for what happens to it.
That has consequences that are easy to state and frequently ignored. A free web tool with no contract is not an acceptable destination for patient data, regardless of how useful it is. Where a tool is contracted, the terms need to say what may be done with the data, whether it is used for training, where it is stored and for how long, and what happens at termination. Under India's data protection framework, purpose limitation applies: data collected for care is not automatically available for training somebody's model.
De-identification helps and is not a magic word. Free-text clinical notes carry identifying detail in ways that automated redaction misses, and small populations re-identify easily. Treat de-identified as reduced risk, not absent risk.
Where a tool makes a claim about clinical purpose, there is also a regulatory question about whether it is a medical device and what approvals apply. Ask the vendor directly and get the answer in writing.
Monitoring, because performance is not a fixed property
A model that performed well on the vendor's validation set may perform differently on your patients, and the only way to know is to measure it on your own.
Drift is normal and has mundane causes: a new scanner, a change in case mix, a different referral pattern, a software update on the model side. So the tool needs a baseline established locally at deployment, a metric that is checked on a schedule, and a person whose job it is to look.
The most informative signal is usually free and rarely collected: what clinicians do with the output. Override and dismissal rates, tracked over time, tell you more than any accuracy figure. A flag that is dismissed almost every time has become noise, and noise trains people to ignore the next flag too. A tool nobody overrides may be trusted, or may be being rubber-stamped, and the two look identical on a dashboard.
What a small hospital can do without a committee
Governance does not require a large programme, and treating it as one is why it never starts.
- Write down every AI tool in use. Ask the departments; the list will be longer than IT's.
- For each, name an owner and write one line on what it may and may not be used for.
- Rule that no patient data goes to any tool without a contract. Give people a compliant alternative in the same breath, or the rule creates shadow usage.
- Ensure high-risk tools log the suggestion, the human's action, and the identity of the human.
- Pick one metric per high-risk tool and review it quarterly, using your own cases.
- Tell patients, in your privacy notice, that AI tools assist in care — and be able to explain how if asked.
The hospitals that get value from AI are not the ones that adopted earliest. They are the ones that can say, a year later, which tools are actually being used, what they changed, and who is answerable — and that can therefore keep the ones that work and drop the ones that do not.


