Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

Can AI help healthcare records answer some of the same questions as clinical trials? A new Oxford study suggests the answer depends less on the algorithms than on the information recorded about patients.

Routinely collected healthcare data are becoming increasingly important for understanding how treatments work in real-world settings. Researchers are exploring whether these records can complement clinical trials by answering some treatment questions more quickly, more efficiently and across larger, more diverse patient populations.

A new study published in Nature Communications, has found that even advanced artificial intelligence (AI) methods could not reliably reproduce clinical trial results when important information was missing from routine healthcare records.

However, when the researchers tested the same methods using simulated data in which all relevant clinical information was available, every approach, including the AI model, correctly recovered the known treatment effects.

Led by DPhil student, Zhengxian Fan, at the Nuffield Department of Women's & Reproductive Health, and supported by senior researchers on the department's Deep Medicine programme, the study reveals that the limitation lies not with AI itself but with the quality and completeness of the available data. When tested using simulated data with complete patient information, the same AI methods successfully recovered the correct treatment effects.

Understanding the challenge

Randomised controlled trials remain the most reliable way of determining whether medicines work. However, they can be expensive, time-consuming and are not always practical. As a result, researchers are increasingly exploring whether routinely collected healthcare records, such as GP and hospital data, can be used to answer similar questions.

One widely adopted approach, known as target trial emulation, aims to analyse observational data in a way that closely mirrors a clinical trial.

While this approach has gained increasing interest, including recognition within NICE's Real-World Evidence Framework, an important question remains: can routine healthcare records reliably reproduce the answers provided by clinical trials?

research method

To investigate this, the research team analysed more than 1.6 million UK patient records from people living with heart failure. They focused on two medicines whose effects are already well established through randomised clinical trials:

  • Beta-blockers, which are known to improve survival.
  • Digoxin, which is known to have a neutral effect on survival.

The researchers then used four different approaches to estimate treatment effects from routine healthcare records, including traditional statistical methods and a Transformer-based deep learning model.

Finally, they repeated the analyses using simulated data in which all relevant clinical information was known.

 

THE FINDINGS

Across the routine healthcare records, none of the four methods reproduced the results seen in the original clinical trials.

Beta-blockers, which are known to improve survival, appeared neutral or even harmful. Digoxin, which clinical trials have shown does not reduce mortality, appeared harmful.

However, when the same methods were tested using simulated data containing complete patient information, all four approaches, including the AI model, successfully recovered the correct treatment effects.

The study showed that improving algorithms alone cannot compensate for information that was never recorded. When researchers tested the same methods using complete patient information in simulations, every approach recovered the correct answer.

 

significance of the results

This study highlights an important limitation of using routine healthcare records to answer questions about treatment effectiveness

Patients who receive different treatments are often different in ways that are not fully captured in healthcare records. Doctors may prescribe medicines based on disease severity, frailty or other aspects of clinical judgement, but these factors are not always recorded. Without that information, even sophisticated statistical and AI methods cannot fully account for differences between patients.

Importantly, the findings do not suggest that AI or real-world healthcare data are unreliable. Instead, they show that the methods work when the data are complete. In the real world, the data are not.

The study reinforces the importance of randomised clinical trials for high-stakes clinical decisions, while also highlighting opportunities to improve routinely collected healthcare data so that future observational studies can produce more reliable evidence.

"We built and tested all types of models that our field is most excited about. In a clean simulation, they worked. On real records, they did not, and the reason is simple: you cannot compute your way to information that was never recorded. This is a data problem, which is exactly why we need robust analyses and continued validation.– Dr Shishir Rao, Supervising Author, AI research lead, Deep Medicine Progamme, Nuffield Department of Women's & Reproductive Health

 

  

Target trial emulation from routine data can improve transparency and design, but it does not remove the deepest problem in observational research. Where the clinical stakes are high and a trial is feasible, we should still run the trial.– Professor Kazem Rahimi, Supervising Author, Director of the Deep Medicine programme the Nuffield Department of Women's & Reproductive Health

Outlook

The researchers believe target trial emulation will continue to play an important role in health research, particularly where clinical trials are impractical or impossible.

Their findings suggest that future progress will depend not only on more sophisticated AI, but also on richer healthcare data and improved recording of the clinical information that influences treatment decisions.

 

Contributions & Collaborators

This research was funded by the British Heart Foundation and the Horizon Europe AI4HF consortium and was led by members of the Deep Medicine programme at Nuffield Department of Women's & Reproductive Health.

The department's research team included Zhengxian Fan (first author), Dr Shishir Rao and Professor Kazem Rahimi (co-corresponding authors), together with Qianqian Yang and YiFan Hu. The study also involved collaborators from the Harvard T.H. Chan School of Public Health and the University of Bristol.

 

Publication

Read the full publication in Nature Communications.

Our Research Groups