Human impact

An AI Wrote a Diagnosis This Patient Never Had

6 min read

Who's checking what the AI writes in your medical record?

The free AI newsletter
An AI Wrote a Diagnosis This Patient Never Had

A woman opens her scan report in her NHS patient record. It says she has demyelination (serious nerve damage that can lead to multiple sclerosis, as the Guardian later put it). The original report said the opposite: "null demyelination." An AI-generated summary had kept the word and dropped the negative.

The woman works for the NHS herself. She is the one who queried the hospital, and the hospital corrected it. She told the Guardian, which broke the story on 31 August: "This was eventually corrected but was a very traumatising experience to be given an incorrect diagnosis because of AI and then be told it's a typo." She asked not to be named.

What Healthwatch collected, and what it doesn't tell us

Healthwatch England is the statutory body that represents patients across the English health system. In July it published research on a category of tools it calls AI scribes: software that listens to the consultation and writes the note that lands in the patient record. Two instruments went into it: a representative YouGov poll of 4,039 adults, and a small opt-in online survey filled in by 44 people.

Those 44 responses are where the error stories come from. Thirty of them involve people who had actually had one of these tools used on them in the past year, and the report says "a small number" found errors in their file. Worth saying upfront: that's not a frequency measure, and Healthwatch doesn't pretend otherwise.

What these accounts do establish is a pattern. Among the cases the Guardian pulled from the file: a drug mixed up with another of a similar name, a discharge letter that fails to mention a prescription needs renewing. The first isn't new: an Ontario audit of twenty AI scribes found the same problem in twelve of them back in May. And in every reported case, the person who caught the error was the patient, not the clinician. Healthwatch says it heard "multiple stories from patients who have noticed these errors when a health professional hasn't."

The sentence the regulator wrote a month earlier

On 29 July, the MHRA, the UK's medicines and medical devices regulator, published guidance settling the status of these tools. Its conclusion: software that only transcribes, summarises, formats, suggests codes, or drafts a letter is not a medical device. That takes it outside the oversight regime that applies to a scanner or a diagnostic aid.

The guidance gives five examples of exempted products. All five carry the same clause, repeated almost word for word: outputs are provided "for clinician review and editing or correction as needed," and, for consultation summaries, "before it is saved to the patient's electronic health record." The regulator's own statement is blunter still: "Clinicians remain responsible for reviewing and verifying AI generated transcripts, summaries and other outputs before they are used in patient care. This responsibility is unchanged by the guidance."

It's a locked door with the key hanging next to it. The exemption doesn't say these tools are safe. It says they don't need regulating like a medical device, because a human checks the output afterwards. The logic holds together perfectly. It rests entirely on a review that nobody measures.

NHS England's own planning framework states that "providers should deploy ambient voice technology (AVT) at pace, with due regard to the national AVT registry," and that trusts should support primary care providers to do so, "ensuring the time freed up is used to see additional patients." That's not an inference. It's written down.

The national supplier registry, meanwhile, is self-certified. NHS England says on its own page that it "does not endorse any of the suppliers on the AVT Supplier Registry" and that it "will only undertake preliminary completion checks against the requirements and standards," leaving the full assessment to whichever trust or practice is buying. As for legal liability, NHS England's guidance calls it "complex and largely uncharted, with limited case law to provide clarity," and warns it can fall on the provider under "a non-delegable duty on the part of the Trust or primary care provider."

The regulator leans on clinician review. The clinician's employer tells them to move faster. The registry checks only at the preliminary stage. And the only verification anyone can point to, in the cases that surfaced, is a patient who opened her own file.

The counter-argument holds up, and it doesn't go away

One thing complicates any straightforward "the AI got it wrong" argument. Charlotte Blease, a researcher at Uppsala University, put it to the Guardian: "The fact is, doctors can and do make mistakes without AI. And it is certainly possible the error rate is worse." In her survey of 1,003 UK GPs, more than half believed their AI-generated records were more accurate than the notes they wrote themselves.

Is there a measurement that settles it? Two recent studies used the same framework to score clinical notes and reached opposite conclusions. One, from the US Veterans Health Administration and published in April in the Annals of Internal Medicine, had eleven tools and eighteen clinicians write notes for five standardised consultations: the human notes won every time. A preprint posted on 18 August, covering 385 simulated consultations across five countries, found the reverse, with critical errors two and a half times more common among junior and mid-level doctors than among the AI.

That second study needs its caveat attached in the same breath: it was funded entirely by Heidi Health, the maker of the tool being evaluated, and every one of its authors is employed by that funder. They disclose it plainly. Heidi Health is listed on the NHS supplier registry.

The number that changes the question

That same preprint, funded by the vendor whose tool it evaluates, contains a finding that cuts both ways. When clinicians review a note with no assistance, they catch about 12% of the errors that a more sensitive method turns up. And they catch an even smaller share of the errors in AI-written notes than in notes written by a colleague.

Which one writes better starts to feel like the wrong question. The safety net the entire regulatory structure leans on (human review) misses most of what it's supposed to catch. Healthwatch puts it more cautiously, without taking sides: it is unclear at this point whether AI scribes lead to more or fewer inaccuracies in summaries compared with clinicians' written notes.

Healthwatch isn't calling for a ban. It wants somewhere for a patient to flag an error, and someone whose job it is to fix it. Its ask of regulators comes down to that: say, finally, who a patient should contact when a note is wrong. In the meantime, 81% of the people surveyed in England want to be asked for consent before the microphone goes on, and nearly nine in ten of those with a recent appointment didn't know whether they had been.

The woman who caught her own false diagnosis did the whole system a favour without meaning to. She was the only reviewer on duty, and her name appears in none of the paperwork.

Topics covered:

HealthAnalysis

Frequently asked questions

What is this AI tool doing during a consultation?
It's software that listens to the consultation and writes the note that goes into the patient's file, what's known in the industry as an AI scribe. It transcribes, summarises and formats. It doesn't make a diagnosis.
Did the MHRA approve these tools?
No. On 29 July 2026, the UK regulator decided that these tools are not medical devices, which takes them outside that regime of oversight. The exemption doesn't say they're safe: it rests on outputs being provided for clinician review, editing or correction, and, for consultation summaries, on being checked before they are saved to the patient's electronic health record.
How often do these AI-generated notes contain errors in NHS records?
Nobody knows yet. The error accounts gathered by Healthwatch England come from an opt-in online survey filled in by just 44 people, 30 of whom had actually had one of these tools used on them: that is not a frequency measure. Healthwatch says it is unclear whether these tools produce more or fewer inaccuracies than notes written by clinicians themselves.
Who is responsible when an AI-written note is wrong?
For checking it, the clinician. The MHRA states that clinicians remain responsible for reviewing and verifying AI generated transcripts, summaries and other outputs before they are used in patient care, and that this responsibility is unchanged by the guidance. Legal liability is a separate question: NHS England's own guidance describes it as complex and largely uncharted, with limited case law to provide clarity, and notes it can fall on the trust or primary care provider. Healthwatch wants regulators to say where a patient should turn when they spot a mistake.
Are patients told when one of these tools is used on them?
Not consistently. In Healthwatch's YouGov survey of 4,039 adults in England, 81% of respondents said they would want to be asked for consent before the microphone is switched on, and nearly nine in ten of those who had a recent appointment did not know whether that had happened.
Katja Liersch

Katja Liersch

Co-founder & Journalist

Katja is a journalist and TV producer. With decades of experience in mainstream media, she brings to Declic Media the perspective of those discovering AI: curious, demanding and pragmatic. She ensures every piece of content truly speaks to everyone.

All articles by Katja →
The free AI newsletter