Comparative evaluation of artificial intelligence–assisted literature search tools for identifying clinically meaningful evidence in cardiology
European Heart Journal - Digital Health

Abstract
The rapid expansion of biomedical literature challenges clinicians’ and researchers’ ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references.
We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88–100],
AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human–AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.
Contributors

Maximiliano Jeanneret Medina
Author

Alexandre Renaud
Author

Maryam Dridi
Author

Valentine Pecriaux
Author

Benoit Lequeux
Author

Stephane Lafitte
Author

Baptiste Maille
Author

Aymeric Menet
Author
You may be interested in



