Comparative evaluation of artificial intelligence–assisted literature search tools for identifying clinically meaningful evidence in cardiology

European Heart Journal - Digital Health

28 July 2026
Organised by: Logo
ESC Journals ARRHYTHMIAS AND DEVICE THERAPY Research Methodology Device Therapy

Abstract

AbstractAims

The rapid expansion of biomedical literature challenges clinicians’ and researchers’ ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references.

Methods and results

We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88–100], P < 0.001) and the highest key-reference overlap (60% [43–68], P < 0.001), whereas Scite performed lowest (20% and 10%, respectively). The tool was the primary determinant of performance (partial R2 = 0.50), whereas prompt formulation had no significant effect. In a pre-specified subanalysis restricted to clinical studies, ChatGPT-5 and human-conducted systematic reviews overlapped by 42% (96% of shared articles highly relevant), with 58% distinct references, indicating complementary AI and human retrieval; ChatGPT-5 produced hallucinated citations when long reference lists were requested for emerging topics, underscoring the need for human verification.

Conclusion

AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human–AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.