Citation

BibTex format

@article{van:2026:10.1136/bmjhci-2025-101959,
author = {van, Kessel R and Anderson, M and McMillan, B and Matthews, MR and Rust, P and Pearcy, P and Nasir, K and Mossialos, E},
doi = {10.1136/bmjhci-2025-101959},
journal = {BMJ Health Care Inform},
title = {Omission and hallucination prevalence of clinical guidelines in diagnostic large language model outputs.},
url = {http://dx.doi.org/10.1136/bmjhci-2025-101959},
volume = {33},
year = {2026}
}

RIS format (EndNote, RefMan)

TY  - JOUR
AB - OBJECTIVE: Meaningful assessments of how large language models (LLMs) incorporate clinical guidelines require large-scale testing over many queries. Here, we evaluate the prevalence of clinical guideline omissions and hallucinations in a large sample of diagnostic LLM outputs. METHODS: We used simulated case vignettes and zero-shot prompting to generate diagnostic outputs and rationales from GPT-4.1 and DeepSeek-V3. English case vignettes were created for hypercholesterolaemia and type-2 diabetes mellitus. Each vignette contained identical medical information, while sociodemographic characteristics varied in terms of sex, ethnicity and location. We calculated the prevalence of existing and hallucinated clinical guidelines in LLM outputs across disease, LLM and sociodemographic characteristics. RESULTS: We analysed a total of 12 197 LLM outputs, which quantifies three hazard areas: omissions (up to 97% for DeepSeek-V3 and 46% for GPT-4.1), hallucinations (up to 9%) and inconsistencies (guideline citation rate ranging from 0% to 78.39% across sociodemographic vignettes). Omission and hallucination rates were generally similar across vignettes with different sex or ethnicity data, yet were particularly sensitive to patient location. DISCUSSION: This study highlights significant variability in clinical guideline prediction across two different diseases, three different sociodemographic variables and two LLMs, even when the LLMs were instructed by identical prompts, establishing clinical guideline prediction in LLM outputs as a stochastic event. CONCLUSION: The stochastic nature of LLMs creates a unique challenge for evidence generation and clinical deployment. Being able to measure and capture this stochasticity within high-quality research designs will be a prerequisite to advancing the responsible deployment of LLMs in healthcare.
AU - van,Kessel R
AU - Anderson,M
AU - McMillan,B
AU - Matthews,MR
AU - Rust,P
AU - Pearcy,P
AU - Nasir,K
AU - Mossialos,E
DO - 10.1136/bmjhci-2025-101959
PY - 2026///
TI - Omission and hallucination prevalence of clinical guidelines in diagnostic large language model outputs.
T2 - BMJ Health Care Inform
UR - http://dx.doi.org/10.1136/bmjhci-2025-101959
UR - https://www.ncbi.nlm.nih.gov/pubmed/42031418
VL - 33
ER -