As ambient AI scribes become routine in general practice, GP, Clinical Safety Officer and Clinical Director for UK & I at Heidi, Dr Tim Cooper says the missing safeguard isn’t regulation – it’s training clinicians to read the output critically.

According to the recent BMJ review around 40% of GPs in the UK now use an ambient scribe. In July, the MHRA clarified that ambient scribes intended solely for transcription, summarising clinical conversations and drafting clinical documentation are not regulated as medical devices under the current framework. Clinician review remains a critical safety control before that output reaches the patient record.

The National Commission into the Regulation of AI in Healthcare has now published its recommendations, and the direction is away from a single approval at launch towards staged authorisation, real-world evidence and continuous monitoring across a product’s life. It is the right direction. It also enlarges the clinician’s job. The safety control for a technology already in daily use across clinical practice is a clinician reading carefully, and almost nobody has been trained to do it.

I should declare an interest. I am a GP and I work for a company that sells this technology, so I have every commercial reason to talk about benefit rather than risk. I am writing about the risk because it is the half of this that suppliers alone cannot fix, and because the opportunity is too good to lose to a failure of preparation.

Training is the missing mitigation in clinical AI

Digital change in healthcare arrives at pace. Video consultations and virtual wards became normal faster than preparation allowed, and we caught up afterwards on the goodwill of clinical teams and the sympathy of our patients. Ambient voice is on the same curve but even faster.

A review this month in BMJ Digital Health & AI, led by researchers at the University of Edinburgh across 27 studies, reflects that ambient scribes capture what is said but can miss what is meant. Facial expression and emotional state, what my psychiatry colleagues would call “affect”, do not always reach the transcript unless vocalised. Summaries risk privileging clinical fact and in primary care it is often the ideas, concerns and expectations that shape the plan.

NHS England’s guidance names an overlapping set of hazards, including transcription inaccuracy, hallucination and uneven recognition of accents and dialects. Patients in Rotherham found the last of these this summer, unable to make themselves understood to their surgery’s AI phone line, and some gave up and walked in instead. None of these are reasons to stop, but every one is a reason to train those wielding the tool.

The system’s own safety infrastructure is sparser than most people assume. A national study in JMIR found that 70% of digital health technologies in NHS trusts and integrated care boards had no documented clinical safety assurance. Organisations reported roughly one full time equivalent Clinical Safety Officer, 1.3 in trusts and 0.4 in integrated care boards, almost always alongside a substantive job. These are not people failing at safety oversight, they are at capacity and doing it in the gaps of a week that was already full.

Our own survey of NHS healthworkers this year found 90% using AI in practice, and 65% doing so before anyone had told them how. I read that as a workforce under pressure reaching for something that helps them get the job done, and on time for once.

I remember learning the hierarchy of evidence as a student. It was drilled into us because a plausible claim and a well founded one look identical on the page. There’s a reason we weight RCTs over observational studies. We have no equivalent taught for the output of a language model. Fluency is not correctness, and nothing in current training teaches a clinician to tell them apart.

So what should training cover? Not the tools, as they will have changed before the module is built. What lasts are the principles of use. Knowing the difference between a general chatbot, a clinical tool built on constrained sources and a regulated device. Reading the note before signing, in a fixed order rather than with a general intention to be careful: the history in the right time frame, the medications right, the plan the one you agreed. Finally, consent, a conversation to be had and recorded rather than a line delivered on the way through the door.

Training is the missing mitigation in clinical AI

The Commission names workforce training as part of the answer. Three things would make it real, and none requires slowing adoption down.

  1. Make appraisal of AI output an assessable competency in Royal College curricula, and name it in GMC standards, so the skill is taught rather than absorbed.
  2. Fund Clinical Safety Officer capacity as a role in its own right rather than a line in an existing overstretched job plan.
  3. Make ongoing training and competency support a procurement requirement, so suppliers compete on how well their tools are used in year three, not on how well they pitch a year-one price.

Suppliers carry real obligations and we should quite rightly be held to them, including for how our tools behave in use long after procurement rather than only at the point of sale. The Commission’s emphasis on responsibility shared between manufacturers, providers and clinicians is welcome, and suppliers should be the first to say so. But a hazard log cannot mitigate a risk that only appears in the room, at the moment a clinician signs a note they have not read. That mitigation is training, and it sits with the system.

I remain optimistic about these tools. Adoption at this speed is a vote of confidence from a profession not easily impressed. The Commission has set a direction and a government response will follow but closing that training gap is the most useful thing policymakers, systems and professional bodies could do this year.

Dr Tim Cooper is a GP and Clinical Safety Officer and Clinical Director UK & Ireland, Heidi.