Echocardiography image quality assessment: human subjectivity and artificial intelligence prediction in multi-centre data

European Heart Journal - Cardiovascular Imaging

29 January 2025
Organised by: Logo
ESC Journals

Abstract

AbstractBackground

Echocardiography image quality is important for precise care, but its assessment is subjective. An image quality metric by artificial intelligence (AI) may permit the optimization of image acquisition and improve analysis.

Purpose

To investigate (i) the most relevant features of echocardiography images related to perceived quality, (ii) to quantify them with AI and (iii) test their variability among centres.

Methods

Fifteen thousand anonymized transthoracic echocardiography studies acquired at Vall d’Hebron Hospital (VH) during care were retrieved. The original clinical assessment of study quality was obtained. Via co-creation, three image quality metrics (each with 3 levels, [0, 1, 2]) were identified: (i) the definition of structures borders and the presence of (ii) all expected cardiac structures and (iii) of apical foreshortening. Then, cardiologists annotated image quality of 6831 VH and 1000 CAMUS videos, which provided its own quality metric [1]. Three AI models were trained with 5256 videos from the VH to reproduce the 3 quality metrics. Models were then internally and externally validated in 1424 and 1000 videos from VH and CAMUS, respectively.

Results

Data from 488 VH patients were available for this analysis (Table 1). Inter-observer agreement in VH data was limited (56, 48 and 46% for border definition, presence of structures and foreshortening, respectively). Borders definition (p<0.001) but not the presence of all expected structures (p=0.862) nor apical foreshortening (p=0.440) was related to clinical study quality.

Trained with data from a single rater, AI models demonstrated reasonable performance in the predictions for borders definition (agreement 60%), presence of all expected structures (54%) and apical foreshortening (53%) in VH images, showing slightly better performance when tested with data annotated by the same rater (62%, 58%, 51%).

Border definition decreased with BMI (p=0.042), which was reproduced by the model (p=0.001), while annotated (p=0.028) but not predicted (p=0.525) border definition was lower in cardiomyopathy, a potential bias of the model. Age, sex, aortic stenosis and irregular heart rhythm did not show trends with border definition.

In CAMUS data, annotated image quality metrics did not compare well with original annotations (agreement of 21% and 35% for border definition and complete structure), which showed lower image quality compared to VH (60 vs 11% of CAMUS videos defined as low quality by VH vs CAMUS raters), highlighting the high subjectivity of image quality. Generalizability to CAMUS images was good for border definition (agreement of 57%) but very limited for structures completeness (19%).

Conclusions

Echocardiography image quality is subjective and mainly driven by the definition of the borders of cardiac structures. AI models can learn image quality, but generalization to images from other centres may be limited.

Demographic and clinical information

Contributors

P Lopez-Gutierrez
P Lopez-Gutierrez

Author

Vall d'Hebron Research Institute (VHIR) Barcelona , Spain

A Morales
A Morales

Author

Vall d'Hebron Research Institute (VHIR) Barcelona , Spain

L Dux-Santoy
L Dux-Santoy

Author

Vall d'Hebron Research Institute (VHIR) Barcelona , Spain

H Majul
H Majul

Author

G Prado
G Prado

Author

G Casas-Masnou
G Casas-Masnou

Author

University Hospital Vall d'Hebron Barcelona , Spain

A Guala
A Guala

Author

University Hospital Vall d'Hebron Barcelona , Spain