Multi-Expert consensus as a label-free quality surrogate for Echocardiographic LV Segmentation

Ufumaka, Isreal, Alibakhshi, Alireza, Fernandes, Patricia, Abdi, Abas, Hussain, Arshian, Dadashiserej, Nasim, Shun-Shin, Matthew, Francis, Darrel and Zolgharni, Massoud ORCID logoORCID: https://orcid.org/0000-0003-0904-2904 (2026) Multi-Expert consensus as a label-free quality surrogate for Echocardiographic LV Segmentation. In: Artificial Intelligence in Healthcare. Lecture Notes in Computer Science, 16877. Springer, pp. 235-247.

Full text not available from this repository.

Abstract

Automated left ventricular (LV) segmentation is fundamental to echocardiographic assessment of cardiac function, yet most models report aggregate performance without accounting for variation in image quality. In clinical practice, poor acoustic windows produce ambiguous endocardial boundaries where both automated predictions and expert annotations become unreliable. We propose a label-free quality surrogate, Qdice, derived from pairwise Dice disagreement across 11 independent clinical experts, requiring no explicit quality labels. Using this surrogate, we conduct a quality-stratified evaluation of three architecturally distinct models (T1, T2, and T3) on the UnityLV-MultiX dataset. T1 is a sparse keypoint model, T2 uses a dense binary mask representation, and T3 is a triple-head hybrid architecture in which predicted keypoint heatmaps guide segmentation feature attention through a differentiable spatial gate. All models are trained on a single-expert dataset and evaluated against a multi-expert consensus using Dice, HD95, MSD, and ejection fraction MAE. A leave-one-out analysis enables direct comparison of model and expert consistency against a common multi-expert reference. All three models exceed every individual expert in overall Dice. T3 achieves the highest Dice across all quality bands (0.937, 0.952, and 0.960 at Qlow , Qmid, and Qhigh respectively; 0.949 overall). It shows the lowest EF MAE overall (5.24%), representing a 38% reduction relative to the expert mean (8.46%). T3’s keypoint and segmentation heads agree far more closely with each other than any cross-model pair (p=0.856), yet residual disagreement between the two heads concentrates on low-quality frames, providing a built-in, single-inference reliability flag without any external reference.

Item Type: Book Chapter or Section
Identifier: 10.1007/978-3-032-35393-1_18
Additional Information: Book chapter of a conference paper presented in the International Conference on AI in Healthcare (AIiH), 28-28 August, London, United Kingdom.
Related URLs:
Date Deposited: 24 Sep 2026
Dates:
Date
Publication status
22 May 2026
Accepted
12 August 2026
Published Online
School, department or research centre: School of Computing and Engineering
THRIVE (The Centre for Translational Healthcare Research, Innovation, Vision, and Excellence)
URI: https://repository.uwl.ac.uk/id/eprint/15318

Actions (admin access)

View Item

Menu