Using natural language processing for automated classification of disease and to identify misclassified ICD codes in cardiac disease
European Heart Journal - Digital Health

Abstract
ICD codes are used for classification of hospitalizations. The codes are used for administrative, financial, and research purposes. It is known, however, that errors occur. Natural language processing (NLP) offers promising solutions for optimizing the process. To investigate methods for automatic classification of disease in unstructured medical records using NLP and to compare these to conventional ICD coding.
Two datasets were used: the open-source Medical Information Mart for Intensive Care (MIMIC)-III dataset (
A newly developed NLP algorithm attained a high accuracy for classifying disease in medical records. XGBoost outperformed the deep learning technique BioBERT. NLP algorithms could be used to identify ICD-coding errors and optimize and support the ICD-coding process.
Contributors

Dries Godderis
Author

Martijn Scherrenberg
Author

Sevda Ece Kizilkilic
Author

Linqi Xu
Author

Marc Mertens
Author

Jan Jansen
Author

Pascal Legroux
Author

Hanne Kindermans
Author

Frank Neven
Author
You may be interested in





