A discharge nurse has eleven patients to move before the end of her shift. Somewhere on her screen, a model has scored one of them an 82.
She sees it. She does not act on it.
Most of this industry treats that as a training problem, a change management problem, or a workflow problem. It is none of those. Clinicians ignore risk scores because the scores have earned it.
We taught them this
In 2021, researchers at Michigan Medicine published an external validation of a sepsis prediction model already running inside one of the most widely used electronic health records in the country. Across 38,455 hospitalizations, its discrimination came in at 0.63. It failed to identify roughly two thirds of the patients who went on to develop sepsis. It generated alerts on 18 percent of everyone admitted, and when it did fire, the chance the patient actually had sepsis was about one in eight.
Sit with what that teaches a clinician. Nearly one in five of your patients trips the alarm, and most of the time the alarm is wrong. After a few months of that, ignoring it is not negligence. It is pattern recognition.
The wider evidence says the same thing. A systematic review in JMIR Medical Informatics covering 23 studies found override rates for decision support alerts ranging from 46 to 96 percent. Nine out of ten alerts dismissed is not a profession resisting technology. It is a profession responding rationally to a signal that rarely pays off.
Readmission prediction has never been exempt. A systematic review of readmission risk models found the nine tested in large US populations had poor discriminative ability, with c statistics between 0.55 and 0.65. The CMS models for heart failure, heart attack, and pneumonia came in at 0.61, 0.63, and 0.63. In one comparison, internal medicine interns predicting readmissions on instinct did about as well as an established model.
If a tool performs roughly as well as a tired resident's hunch, it has not earned the right to change anybody's plan.
Accuracy is not where these tools fail
The reflex response to all of this is a promise of a better model. More data, newer methods, a higher number on a validation set.
That response misses where the failure actually happens. Even a genuinely accurate score fails in the same place, because the number arrives with nothing attached to it. No reasoning. No recommended action. No owner.
An 82 tells a nurse nothing she can use. It is a verdict with no argument behind it, handed down by something she has no way to question. Her options are to trust it blindly or disregard it, and no experienced clinician is going to choose the first one.
Now change what appears on her screen. This patient's risk is elevated because four medications changed at discharge, there have been two emergency visits in the past six months, and no follow up is scheduled.
She is no longer deciding whether to trust a model. She is checking whether three specific things are true about her patient. That is a question clinicians answer all day, in seconds. And every one of those three facts points at something a person can actually do before the patient leaves the building.
Reasoning does something more useful than build trust
It lets the clinician disagree in a way that helps.
A bare score can only be accepted or ignored, and those are the only two things anyone ever does with one. Reasoning can be corrected. A nurse who sees that the flag rests partly on a missing follow up can tell you the appointment exists and simply has not made it into the system yet. That correction improves the next signal. A naked number gives her nowhere to put what she knows, and what she knows is usually the most current information in the room.
This is the part the market still has backwards. Explainability gets treated as a compliance checkbox or a nice extra to mention on the third slide. It is neither. It is the line between a signal a clinician can work with and a signal a clinician learns to scroll past.
The same failure, a different setting
This is not a discharge problem. It shows up anywhere a score gets handed to a busy person.
Telling a care team a patient is likely to stop GLP 1 therapy is useless. Telling them the patient has not filled a prescription in forty days and has not been seen since therapy started is a phone call somebody can make this afternoon.
Two different problems, one rule. A score describes. Reasoning directs.
Better questions to ask
If you are evaluating anything that produces a risk signal, stop opening with how accurate it is. Accuracy is table stakes and it is not the bottleneck. Three questions matter more.
Can it show me, for this specific patient, what is driving the score?
Can my team act on it in under a minute, without leaving the work they are already doing?
Can we tell afterward whether the action made any difference?
Anything that cannot do all three will end up ignored no matter how well it performs on paper. Your team will teach itself to scroll past it, and they will be right to.
The measure of a risk score was never how often it is correct. It is how often somebody does something different because of it.
At Preventra, every score carries its reasoning at the patient level, because a number nobody acts on is not intelligence. It is noise with a decimal point.
Sources
- Wong A, et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 2021. View study
- Poly TN, et al. Appropriateness of Overridden Alerts in Computerized Physician Order Entry: Systematic Review. JMIR Medical Informatics, 2020. View study
- Kansagara D, et al. Risk Prediction Models for Hospital Readmission: A Systematic Review. JAMA, 2011. View study