Falls data: why building the dataset comes before building the model

by Laura Bassett

Dr Ekin Yağış and Joy Li

Falls are among the most common patient safety incidents in healthcare. They can prolong hospital stays, delay recovery, and place additional pressure on already stretched services. While identifying patients at risk of falling could help support prevention efforts, developing a predictive model is only part of the story. Before any model can be built, there is a less visible but equally important challenge: understanding the data itself.

At iCARE, researchers are exploring how routinely collected healthcare data can support safer, more proactive care. Yet transforming clinical data into something suitable for analysis is rarely straightforward. Before a model can identify patterns, researchers must first understand what information exists, how consistently it is recorded, how reliable it is, and whether it accurately reflects the realities of clinical practice.   

To better understand that process, I spoke with iCARE's data scientists, Joy Li and Dr Ekin Yağış, who are researching inpatient falls prevention. Joy trained in information engineering before moving into biomedical data science, with research experience at Bosch and Cambridge. Ekin holds a PhD in deep learning for disease detection and is now a Research Associate at iCARE. Their reflections offer insight into a stage of data science that often receives little attention but can have a significant impact on the quality, reliability, and usefulness of any future analysis. 

Records of care, not measurements for research 

Hospital data exist because patients are being looked after. That seems obvious, but its consequences are easily underestimated. "Most clinical data are collected as part of the patients' healthcare journey and to support their care rather than research," Joy explains, "so there are many steps between getting access to the data and being able to analyse them." 

Ekin argues that this undermines a common proxy for data quality. "People often assume that if a field is filled in, it's reliable," she says. "In reality, some of the most complete datasets are the least useful for research because they were recorded for administrative rather than clinical purposes. Conversely, sparse but carefully documented data from engaged clinicians can be far more valuable." Completeness describes how often something was recorded. It says nothing about why, by whom, or how faithfully it reflects what happened. 

Making sense of fragmented data 

When looking at the final dataset, it can be easy to forget how imperfect and fragmented the raw data were. A fall is not a single data point, but an event surrounded by information recorded before, during and after it, often in different parts of the medical record. 

We found that falls involve pre-fall assessments, records during the patient's admission and post-fall evaluations, all appearing in different parts of the medical record,” Joy explains. “One of our first challenges was not modelling the data, but bringing these different data sources together and understanding their relationships. 

That process, known as data curation, involves much more than simply cleaning a dataset. It means making decisions about what information to include, exclude or transform, and understanding what different fields actually represent. 

All of the curation is hidden in the final dataset,” Ekin says. “The hundreds of decisions about what to include, exclude or transform. The conversations with clinicians about what fields actually mean. The instances where we had to exclude data because it was ambiguous or unreliable.” 

Behind those decisions are questions about what happened in clinical practice. Why was something recorded in one place but not another? Does a missing value mean something did not happen, or simply that it was not documented? Answering those questions means looking beyond the dataset and understanding the clinical environment in which it was created. 

For Ekin, that means working closely with the teams who know that environment best. "I learned that the best data insights require embedding yourself in the clinical environment, understanding why things are recorded the way they are, and building trust with the teams who could actually use the work," she adds. That clinical insight can help make sense of differences in recording practices, including instances where falls were coded differently between wards. 

For the iCARE team, that clinical context is not separate from the data science. It is part of what makes meaningful analysis possible. 

By bringing together information from across the patient journey and understanding the context in which it was recorded, researchers can begin to build a more reliable picture of falls and identify patterns that might otherwise remain hidden. 

Looking beyond the data 

With a more robust, better-understood dataset in place, the team can begin to explore the differences between falls, recognising that not all falls are the same and that prevention should not be one-size-fits-all. The aim is not simply to find patterns in the data, but to understand which of those patterns could ultimately be useful in supporting safer care. 

There are lots of things that excite me about this work,” Joy says, and one of them is how data science could potentially have an impact on patients, clinical care, and even healthcare policy. 

Falls aren't just data points for a risk model; they're moments of fear, injury, loss of confidence, sometimes permanent harm. This realisation keeps me grounded about the responsibility we have to get this right. Dr Ekin Yağış Data Scientist at iCARE

But with that potential comes responsibility. For Ekin, keeping sight of the patient behind the data is fundamental. 

Every row in a dataset is a person, an event, often a moment where something went wrong,” she says. “Falls aren't just data points for a risk model; they're moments of fear, injury, loss of confidence, sometimes permanent harm. This realisation keeps me grounded about the responsibility we have to get this right. 

Read what the analysis found [here] 

 

Article text (excluding photos or graphics) © Imperial College London.

Photos and graphics subject to third party copyright used with permission or © Imperial College London.

Article people, mentions and related links

Reporters

Laura Bassett

Faculty of Medicine