Citation

BibTex format

@inbook{Ufumaka:2027:10.1007/978-3-032-35393-1_18,
author = {Ufumaka, I and Alibakhshi, A and Fernandes, P and Abdi, A and Hussain, A and Dadashiserej, N and Shun-Shin, M and Francis, D and Zolgharni, M},
doi = {10.1007/978-3-032-35393-1_18},
pages = {235--247},
title = {Multi-Expert Consensus as a Label-Free Quality Surrogate for Echocardiographic LV Segmentation},
url = {http://dx.doi.org/10.1007/978-3-032-35393-1_18},
year = {2027}
}

RIS format (EndNote, RefMan)

TY  - CHAP
AB - Automated left ventricular (LV) segmentation is fundamental to echocardiographic assessment of cardiac function, yet most models report aggregate performance without accounting for variation in image quality. In clinical practice, poor acoustic windows produce ambiguous endocardial boundaries where both automated predictions and expert annotations become unreliable. We propose a label-free quality surrogate, q<inf>dice</inf>, derived from pairwise Dice disagreement across 11 independent clinical experts, requiring no explicit quality labels. Using this surrogate, we conduct a quality-stratified evaluation of three architecturally distinct models (T1, T2, and T3) on the UnityLV-MultiX dataset. T1 is a sparse keypoint model, T2 uses a dense binary mask representation, and T3 is a triple-head hybrid architecture in which predicted keypoint heatmaps guide segmentation feature attention through a differentiable spatial gate. All models are trained on a single-expert dataset and evaluated against a multi-expert consensus using Dice, HD95, MSD, and ejection fraction MAE. A leave-one-out analysis enables direct comparison of model and expert consistency against a common multi-expert reference. All three models exceed every individual expert in overall Dice. T3 achieves the highest Dice across all quality bands (0.937, 0.952, and 0.960 at Q<inf>low</inf>, Q<inf>mid</inf>, and Q<inf>high</inf> respectively; 0.949 overall). It shows the lowest EF MAE overall (5.24%), representing a 38% reduction relative to the expert mean (8.46%). T3’s keypoint and segmentation heads agree far more closely with each other than any cross-model pair (ρ=0.856), yet residual disagreement between the two heads concentrates on low-quality frames, providing a built-in, single-inference reliability flag without any external reference.
AU - Ufumaka,I
AU - Alibakhshi,A
AU - Fernandes,P
AU - Abdi,A
AU - Hussain,A
AU - Dadashiserej,N
AU - Shun-Shin,M
AU - Francis,D
AU - Zolgharni,M
DO - 10.1007/978-3-032-35393-1_18
EP - 247
PY - 2027///
SP - 235
TI - Multi-Expert Consensus as a Label-Free Quality Surrogate for Echocardiographic LV Segmentation
UR - http://dx.doi.org/10.1007/978-3-032-35393-1_18
ER -