Application of machine learning techniques to optimise for limited datasets in an AI-ECG model for Brugada classification
EP Europace Journal

Abstract
Brugada syndrome is rare but important to diagnose due to its predisposition to sudden death. Deep neural networks can be trained to automate the recognition of the diagnostic type 1 Brugada ECG pattern on a 12-lead ECG with variable performance reported previously. The scarcity of high quality training data hinders the development of robust and accurate AI ECG models in the case of Brugada recognition as well as for other rare cardiac conditions.
We tested the effectiveness of various representation learning approaches for pretraining and oversampling techniques to enhance the performance of a Brugada classification AI-ECG model, particularly with limited available data.
A total of 704 12-lead ECGs were extracted from a hospital database, consisting of 176 ECGs of known Brugada patients, 176 right bundle branch block (RBBB) "Brugada mimic" ECGs and 352 normal ECGs. The pooled data was partitioned so that 80% was used for training and validation of a Brugada classification model and the remaining 20% comprised an unseen test dataset.
Pretraining was undertaken using either supervised techniques or state-of-the-art self-supervised representation learning approaches (SimCLR and MoCo-V2 frameworks). The pretraining ECG data comprised 19648 non-Brugada ECGs (RBBB and normal ECGs) also extracted from the hospital database but not included in the final training or test datasets. The knowledge from the pretrained models formed the basis of separate fine-tuned models for Brugada classification. The synthetic minority oversampling technique (SMOTE) was incorporated during fine-tuning to synthetically enhance available data for training.
All Brugada classification models were evaluated on the unseen test dataset. The models were also trained using lower proportions of the available training data to compare impact on performance with increased training data scarcity.
The non-pretrained baseline model accuracy was 94.5% (F1-score 0.892, AUC 0.984). The best performing pretraining approach utilised the SimCLR framework with SMOTE oversampling, augmenting model accuracy to 97.6% (F1-score 0.949, AUC 0.998), as shown in Table 1.
When lower proportions of training data were available for model development, the supervised pretraining approach demonstrated the best performance, however the self-supervised SimCLR and MoCo-V2 models showed incremental improvements over the baseline Brugada classification model (Panels C & D, Figure 1).
Advanced representation learning techniques provide a substantial added benefit in the performance of Brugada AI-ECG classification models. These learnings can be further applied in developing neural networks for rare cardiac conditions where ECG training data is often limited.
You may be interested in




