Reducing administrative burden in cardio-oncology: an LLM-powered approach for automated summarization of oncology and cardiotoxicity histories in dutch medical records
European Heart Journal Supplements

Abstract
Clinical documentation is a significant administrative burden in cardio-oncology, where comprehensive medical histories—including prior cancer diagnoses, treatments, and cardiotoxicity—are essential for patient care. Large language models (LLMs) using retrieval-augmented generation (RAG) have demonstrated the potential to partly automate this task. However, their performance in Dutch medical texts—and most non-English texts—remains unexplored.
This study evaluates the feasibility, accuracy, and efficiency of a locally deployed, open-source RAG-LLM infrastructure for automatically summarizing oncology and cardiotoxicity histories in Dutch medical records.
At the LUMC cardio-oncology outpatient clinic, 54 consecutive patients were enrolled in this proof-of-concept study. All relevant in- and outpatient letters were retrieved from the institutional data hub. Two investigators manually annotated prior cancer diagnoses, cancer treatments, and cardiotoxicity history, establishing a gold standard dataset. A RAG pipeline was developed and deployed on the institutional high-performance computing cluster using exclusively open-source components. RAG enhances LLM responses by retrieving only relevant text from the EHRs, reducing hallucinations. The pipeline incorporated LLaMA3.1 8B, one of the highest-performing smaller-size LLMs. The system automatically extracted the same key variables from patient records. Three investigators independently evaluated the generated summaries, comparing them to the gold standard for accuracy.
A total of 54 patient records were analyzed, containing 412 documents. The most common prior cancer diagnoses were lymphoma (23/54, 42.6%) and leukemia (13/54, 24.1%). Nearly all patients underwent chemotherapy, with 27 (50.0%) receiving radiotherapy and 16 (29.6%) undergoing stem cell transplantation. The system correctly identified 52/54 cancer diagnoses, 53/54 chemotherapy status, 49/54 radiotherapy status, and 52/54 stem cell transplantations. Systemic antineoplastic agents were fully or for majority identified in 48/54 cases (88.9%). For cardiotoxicity detection, the model correctly identified 9/12 cases and 38/41 non-cases, yielding a sensitivity of 75.0% and a specificity of 92.7%. The LLM demonstrated excellent comprehension of Dutch medical jargon, with all errors stemming from retrieval limitations rather than incorrect reasoning by the model. Ineffective retrieval also contributed to missed cancer treatments. The model completed summaries in an average of 11 seconds per patient, compared to 3.5 minutes for manual extraction.
A locally deployed RAG-LLM infrastructure can accurately and efficiently summarize medical histories in a Dutch cardio-oncology setting. While the system showed high accuracy in cancer diagnosis and treatment extraction, retrieval inefficiencies limited sensitivity for cardiotoxicity detection and complete treatment identification. central illustration
Contributors

M Fatah
Author

C J J Westermann
Author

G Liao
Author

M L Hamer
Author

M M Van Buchem
Author

J W Jukema
Author
You may be interested in




