Evaluation of AI-based ECG analysis accuracy using Chat GPT-4 and Gemini AI models

EP Europace Journal

23 May 2025
Organised by: Logo
ESC Journals

Abstract

AbstractBackground

The application of artificial intelligence (AI) in medicine, including cardiology, is rapidly advancing. Large language models have introduced new possibilities for ECG analysis, though their reliability remains uncertain. This study evaluated the accuracy of ECG interpretations by ChatGPT-4, developed by OpenAI, and Gemini, developed by Google AI.

Methods

We uploaded anonymized ECGs from 130 patients referred for arrhythmology consultation for analysis by the AIs. The control group consisted of a consensus from three experienced electrophysiologist.

Results

Both Chat GPT-4 and Gemini exhibited low accuracy compared to the expert consensus: the most likely diagnoses matched the expert consensus in only 31.21% of cases for Chat GPT-4 and 25.64% for Gemini. Chat GPT-4 performed best in recognizing sinus rhythm, achieving an accuracy of 85.72%, while Gemini identified sinus rhythm correctly in only 28.57% of cases. For Gemini, the best results were observed in the identification of atrial fibrillation, with a 48.89% match rate, compared to a 35.56% accuracy for Chat GPT-4.

Conclusion

The latest Chat GPT-4 and Gemini AI models perform poorly in interpreting ECGs of patients referred to an arrhythmia clinic. Their use currently does not provide adequate assistance and does not replace expert evaluation.

Contributors