Nargiz Noimann-Zander, founder of X-Technology, explains why healthcare organisations need to govern psychological and cognitive inferences as carefully as the raw data used to generate them.
Healthcare AI is learning to do more than record what a patient says or measure what their body does. Systems can analyse voice, gaze, movement, response time and physiological signals, then infer distress, fatigue, pain, confusion, declining attention or cognitive vulnerability.
These capabilities may help clinicians notice changes that are easy to miss during a brief appointment. They can also place an unverified interpretation inside a care pathway with the appearance of objective fact.
A heart rate is a measurement. The statement that a patient is anxious is an inference. A pause before answering is observable. The conclusion that a patient is cognitively impaired requires clinical judgement. Yet these stages can be compressed into a single label on a dashboard or record.
That compression creates a clinical risk. Once a conclusion enters the record, it may influence triage, referral, treatment and the way later clinicians interpret the patient. The organisation therefore needs rules for the conclusion, not only controls around the sensor that produced the original data.
Separate the signal from the conclusion
Voice, gaze and movement are highly dependent on context. Slower speech may reflect anxiety, pain, medication, tiredness, a language difference or an unfamiliar interface. Reduced eye contact may be linked to distress, but it may also reflect culture, vision, discomfort with the camera or a deliberate effort to concentrate. A raised heart rate can accompany fear, exertion, fever or a cup of coffee.
Pattern recognition can identify a correlation. It cannot establish clinical meaning on its own. Providers should therefore preserve three distinct layers wherever an inference appears.
The first layer is the observed data, such as a change in speech rate or response time. The second is the model output, including what the system inferred, which version of the model generated it and how uncertain the output is. The third is the clinical assessment, which records whether a qualified professional accepted, rejected or revised the inference after considering the patient and the circumstances.
If these layers are combined, later users may be unable to tell what the patient actually did, what the system predicted and what a clinician concluded. Provenance becomes part of patient safety.
Precision can disguise uncertainty
An output can look authoritative because it arrives as a score, colour or percentage. That presentation should not be confused with clinical certainty. A confidence score may describe the model’s calculation, but it does not show whether the model was validated for this patient, this setting or this purpose.
The Information Commissioner’s Office guidance on AI accuracy says organisations using AI to make inferences about people must consider the possibility that an inference is wrong and the effect it may have on later decisions. It also advises that records distinguish statistically informed guesses from facts and retain information about the provenance of the inference.
For healthcare providers, overall accuracy is rarely enough. They need to understand false positive and false negative rates, how performance changes across relevant patient groups, and what happens when the environment differs from the conditions used for validation. A system trained on controlled speech samples may behave differently in a noisy ward, during a remote consultation or with someone who has a speech impairment.
The acceptable error balance also depends on the consequence. A low-impact prompt asking a clinician to check in with a patient carries a different risk from a flag that changes access to treatment or triggers an urgent mental health referral.
Inferred data needs explicit governance
Patients may understand that a consultation is being recorded without realising that their voice will also be analysed for signs of anxiety or cognitive change. They may agree to share heart rate data for rehabilitation without expecting it to support a psychological profile.
As current ICO guidance on special category data explains, an organisation that intentionally infers a person’s health status or health risk is processing special category data, regardless of how confident the inference is. The sensitivity lies in the intended conclusion and how it is used, not merely in the raw input.
This matters for procurement and deployment. A provider should know which inferences a system creates, where they are stored, who can see them, how long they remain available and whether they are reused for another purpose. It should also decide whether the inference belongs in the clinical record at all. Some outputs may be appropriate as temporary decision support and unsafe as permanent statements about the patient.
Human oversight has little value if the reviewer sees only the label and is expected to approve it quickly. A meaningful review requires access to the underlying signal, relevant limitations, alternative explanations and the threshold that caused the system to act.
The reviewer also needs training, time and authority to disagree. The ICO’s guidance on individual rights in AI systems emphasises system design that supports human review, appropriate training and the ability for staff to override or escalate an automated decision.
Patients need a route to question the conclusion as well. A person should be able to explain that a voice change followed surgery, that a pause reflected translation or that a movement pattern was caused by pain. That information should travel with the disputed inference so later decisions do not continue to rely on a label that has already been challenged.

Build an inference standard into clinical safety
The UK already has a foundation for this work. DCB0129 covers clinical risk management by manufacturers of health IT, while DCB0160 covers the organisations that deploy and use it. NHS England’s current review of both standards recognises how much digital care has changed and is examining how the framework should develop.
Healthcare organisations do not need to wait for a new national rule before setting a local standard for inferred information. At minimum, that standard should require six things. First, keep observed data, model inference and clinical assessment visibly separate. Then, define the permitted purpose and consequence of every inference before deployment.
Present uncertainty, known limitations and plausible alternative explanations to the reviewer, and require qualified human review before an inference changes care or enters the permanent record. Next, give patients a clear way to question or correct an inference and preserve that challenge in the record.
And finally, monitor errors, overrides, complaints and performance across patient groups after deployment.
These controls should be proportionate to the consequence. A gentle prompt to ask one more question does not require the same assurance as an automated risk classification. The governing question is how far the conclusion can travel and what it can change.
AI can widen clinical attention by detecting patterns that humans may not see. Its value depends on preserving the difference between a useful signal and a verified finding. When uncertainty is visible, review is meaningful, and patients can challenge what has been inferred about them, innovation becomes easier to trust and safer to use.



