Citation

BibTex format

@article{Lee:2026,
author = {Lee, K and Jing, P and Zhang, Z and Yang, Y and Wang, T and Marshall, DC and Fang, Y and Yang, G},
journal = {npj Artificial Intelligence},
title = {Seeing through experts' eyes: a foundational vision-language model trained on radiologists' gaze and reasoning},
year = {2026}
}

RIS format (EndNote, RefMan)

TY  - JOUR
AB - Large-scale vision–language models have shown promise in automating chest X-ray interpretation. However, their clinical utility remains limited by a fundamental gap between model outputs and the reasoning processes of radiologists. Most systems optimize for semantic information without emulating how experts visually examine and interpret medical images. As a result, current models often overlook critical findings, misrepresent anatomical context, or generate reports that diverge from established diagnostic workflows. Radiologists typically follow structured protocols (e.g., the ABCDEF approach that sequentially assesses the airways, breathing, circulation, diaphragm, external and foreign material). This standardized workflow ensures that all clinically relevant regions are examined in a consistent and systematic manner. It reduces the risk of missed findings, supports reliable diagnostic reasoning and facilitates clear communication between radiologists and referring clinicians. Emulating such expert strategies is essential for improving the trustworthiness and interpretability of automated report generation systems. Therefore, we introduce Gaze-X, a vision–language model that leverages radiologists’ eye-tracking data as a behavioral prior to model expert diagnostic reasoning. By incorporating gaze trajectories and fixation patterns into pretraining, Gaze-X learns to follow the spatial and temporal structure of radiologist attention and integratesvisual observations in a clinically meaningful sequence. This approach enables the model to align itsfocus with diagnostically relevant regions and to emulate the interpretive logic underlying expert reports. Using a curated dataset of gaze recordings, comprising over 30,000 key frames across diversedisease categories, from five radiologists interpreting chest X-rays, we demonstrate that Gaze-X produces more accurate, interpretable, and expert-consistent outputs across a range of clinically relevanttasks a
AU - Lee,K
AU - Jing,P
AU - Zhang,Z
AU - Yang,Y
AU - Wang,T
AU - Marshall,DC
AU - Fang,Y
AU - Yang,G
PY - 2026///
SN - 3005-1460
TI - Seeing through experts' eyes: a foundational vision-language model trained on radiologists' gaze and reasoning
T2 - npj Artificial Intelligence
ER -

Contact


For enquiries about the MRI Physics Collective, please contact:

Mary Finnegan
Senior MR Physicist at the Imperial College Healthcare NHS Trust

Pete Lally
Assistant Professor in Magnetic Resonance (MR) Physics at Imperial College

Jan Sedlacik
MR Physicist at the Robert Steiner MR Unit, Hammersmith Hospital Campus