From AI hype to measurable clinical impact: insights from global healthcare experts
30 July 2026
At HLTH Europe 2026, Elsevier convened an invitation-only executive breakfast roundtable bringing together senior healthcare leaders, clinicians, policy experts and technology innovators from across Europe. The session was hosted by Vittorio Sportellio, Sales Director at Elsevier, with opening remarks from Hugh Thomas, Editor-in-Chief of The Lancet Digital Healthopens in new tab/window.
Participants included clinicians, digital health leaders, academics, executives, and policy advisors from across the UK, Germany, Switzerland, the Netherlands and beyond, united by their roles in shaping how AI is adopted in practice.
Rather than a formal presentation, the discussion unfolded as a working conversation shaped by the experiences, evidence and challenges raised directly from the room. At the center of the dialogue was a shared challenge:
AI adoption is outpacing evidence in healthcare. That gap is not closing. The question is how to lead responsibly inside it.
The evidence is harder than the hype
Hugh Thomas opened by grounding the group in two recent landmark studies, both had something important to say about the distance between what AI promises in principle and what it delivers in practice.
The first was the Tricorder trial, published in The Lancet in February 2026: a cluster randomized trial of an AI-enabled stethoscope across 205 GP practices and approximately 1.5 million patients. The headline finding was sobering: no meaningful increase in cardiovascular disease diagnosis at the population level. The per-protocol analysis showed real benefit among the ~12,500 patients where AI was used, but that was the problem, most practices barely used it. Five GP practices contributed 35% of all readings. The issue was not the device. It was workflow friction, specifically the absence of direct EHR integration. It was a theme that recurred throughout the morning, with several of the studies and examples discussed pointing to workflow integration, rather than raw algorithmic performance, as the decisive factor in whether an AI tool is adopted and used in practice.
The second was an observational study on de-skilling: endoscopists exposed to AI during colonoscopy showed a 6-percentage point drop in adenoma detection rate, from ~30% to ~24%, in their unassisted procedures. The effect held across centers and experience levels.
He also presented two studies on algorithmic bias. GPT-4, when fed case vignettes with varying patient race and gender, altered its clinical management recommendations, a finding described as unsurprising to anyone who had looked, but largely unlooked-for. A separate study from Dr. Judy Gichoya’s group showed that deep learning models trained on CT, X-ray and mammography images could predict patient race with an AUC of ~0.9, even when images were degraded to near-unrecognizable levels.
The implication was clear: the evidence base for AI in healthcare is genuinely mixed. Some of the most rigorous, large-scale trials are returning null or cautionary findings. And yet adoption is accelerating.
Measuring what matters
After a show of hands revealed approximately half the room was already using at least one AI tool in clinical practice, Vittorio Sportelli asked a pointed question: how do you count the value? The response exposed a fundamental tension between financial return and human benefit that the field has not resolved.
The cost reality was put plainly by one participant, quality improvement is not in dispute: AI detects cancers that would otherwise be missed. But the efficiency gains that would justify investment at scale have not materialized. Annual AI expenditure in radiology runs from €20,000 for a single application at a small center to €1-2 million for a large multi-system deployment.
There is absolutely no doubt that, if you are using AI as a radiologist, you will improve the quality of your diagnostic work. However, we have not yet achieved the efficiency gains that the institutions paying for these solutions often require. If you’re going to invest one or two million dollars, at some point you need to see a return on that investment.Dr. Amine Korchi
Radiologist
Some participants offered concrete numbers. Automated blood result filing reduced cost per result from 50 pence to 20 pence in one NHS organization, with improved staff retention as a further dividend. One participant raised that a no-show prediction algorithm at their organization increased patient attendance substantially. Aahuti Rai, Partner at Four Points Health, identified a structural problem: organizations piloting AI in isolation accumulate fragmented subscriptions (one institution surveyed had 24 active SaaS contracts) without capturing system-level efficiency. Where ROI exists, it requires platform thinking, not point-solution logic.
Staff satisfaction, reduced burnout, and retention emerged as the metrics participants found more meaningful than financial modelling.
The pivot: from ‘prove’ to ‘implement and mitigate’
The session's most significant exchange came around a question Vittorio Sportelli posed directly: what evidence threshold is good enough to deploy AI clinically? The room's answer, across roles and systems, was not a number or a study design. Instead, a new logic emerged: Deploy. Monitor. Adjust.
A Chief Clinical Information Officer who had taken a class 2B-regulated tool through both clinical trial and real-world deployment described the gap between the two:
If I were to wait for every single startup to go through rigorous academic evidence, I will never deploy. And I have a big problem in my hospital that I need to fix.Dr. Matea Deliu
Chief Clinical Information Officer, Bromley Healthcare/Guy’s and St Thomas’ NHS Foundation Trust
Another participant described using hazard logs, clinical risk categorization tools, as the practical substitute for regulatory standards that don't yet exist.
There are no written clinical standards, so we feel very nervous about just saying: switch off the manual checking and go with the digital tools.Dr. Mina Gupta
Group Clinical Chair, Modality Partnership
This reflects a practical reality. Healthcare systems face immediate pressures—workforce constraints, rising demand, and operational inefficiencies. Waiting for perfect evidence is often not an option.
Instead, organisations are implementing where there is perceived benefit, monitoring for harm, co-developing with vendors in real time, and building the evidence iteratively rather than prospectively. It is a pragmatic response to a genuine gap.
Who owns the risk when you implement and mitigate?
If 'implement and mitigate' is the operating logic, accountability cannot be informal. The session surfaced a governance gap that several participants recognized as one of the more urgent problems in the field, clinical AI deployment is outpacing the structures designed to oversee it.
In many organizations, responsibility for AI has landed by default rather than design, with whoever was willing to pick it up. A medical director from the NHS described why so few clinical leaders have been willing:
People are afraid of digital - that's why I've been doing it for so long. Nobody else wants to get involved, because they're afraid it goes wrong and they're going to take the blame.Dr. Masood Nazir
Medical Director and CCIO
This aversion has consequences that compound over time. When no one owns it, vendors face liability uncertainty and become cautious about EHR integration. When tools sit outside the core workflow, adoption falls. When adoption falls, the perception anchors that are driving deployment never form. The system stalls — not because the technology failed, but because the governance structure around it did.
The group identified two things that need to change. First, clinical safety must become a core competency for the next generation of clinicians, and secondly clinical leadership must move upstream, and not remain a procurement or IT decision.
You will have to learn clinical safety. No clinician should be trained without considering the safety of technology, the impact it would have on patients. You're going to have to say: is this product safe? Is this going to cause harm to our patients?Dr. Masood Nazir
Medical Director and CCIO
The equity gap that 'implement and mitigate' most often misses
The implement-and-mitigate posture carries one risk above all others: the populations most likely to be harmed by biased or poorly validated AI are also the least likely to be represented in the feedback loops that would surface the problem.
These biases are not theoretical. The studies cited at the start of the session on GPT-4’s race- and gender-sensitive recommendations and AI’s ability to predict patient race from degraded medical images, showed that bias is already present in deployed tools, and that too few people are actively looking for it.
One participant brought this to ground level. Her trust serves a city with no ethnic majority, with 100 languages spoken within a mile of the main hospital and many of the communities aren't included in historical datasets on which most deployed tools were trained.
The first conversation that we have to have [with vendors] is about the equity and the bias principle, because otherwise we're building tools that don't work for our whole population.Dr Ruw Abeyratne
Director of Health Equality and Inclusion, University Hospitals of Leicester NHS Trust
A practical example made this tangible. An AI translation tool, initially promising, failed in community testing because its language was too formal and culturally misaligned. It would have been deployed widely had local feedback not intervened early.
Equity, therefore, cannot be retrospective. It must be built into deployment from the outset.
Looking ahead: cautious optimism, shared accountability
The questions raised in this session cannot be answered in ninety minutes, and clinical AI deployment is not waiting for them to be resolved. Governance, standards and equity safeguards are all struggling to keep pace with what is already being deployed. That is not a reason to slow down. It is a reason to build more deliberately.
Three shifts will define whether the next phase delivers on its promise. The first is from perception to evidence: outcome data, bias monitoring, and longitudinal evidence from real clinical environments need to run alongside deployment, not follow it.
The second is from individual effort to institutional accountability. Much of the progress described in the session has depended on individuals taking risks and organizations developing their own frameworks. But shared standards and regulatory models that reflect the adaptive nature of clinical AI are now overdue. Until they exist, individual Trusts are doing the work regulation has not yet caught up with. Whether AI can scale will depend on two things: whether the technology integrates and performs reliably in real-world settings, and whether its clinical outputs are credible enough for clinicians to act on.
The third is from deployment to co-design: the communities most affected by these tools have expectations and the capacity to improve products before they scale. Meeting them early is a requirement, not a courtesy.
The executives in this room are making decisions now, about governance, equity, and what counts as evidence, that will shape a healthcare system substantially defined by AI within a decade. The foundations being laid today will determine whether that system is one clinicians trust, patients benefit from, and institutions can stand behind.