Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Preprocessing of unstructured medical data: the impact of each preprocessing stage on classification

2020·25 Zitationen·Procedia Computer ScienceOpen Access

Volltext beim Verlag öffnen

Zitationen

Autoren

2020

Jahr

Abstract

Nowadays, it is still important to develop methods for processing data, in particular medical texts, in Russian. In this paper, we checked how each stage of text pre-processing affects the result of the classifier. The paper analyzed 269923 records of allergic anamnesis of patients, 11670 of which were placed for further processing. We consider the main stages of pre-processing: tokenization, deletion of stop words, error correction, document cropping, normalization, class harmonization, and vectorization. To vectorize the data, we have selected the Bag-of-Words. The method of logistic regression was chosen for classification, since it has easy reproducibility and interpretation. Precision, recall and F-measure were selected as evaluation metrics. The results (F = 88.12%) showed that the most effective was the stage of normalization and error correction.

Autoren

Institutionen

ITMO University(RU)

Themen

Machine Learning in HealthcareTopic ModelingBiomedical Text Mining and Ontologies

Volltext beim Verlag öffnen

Preprocessing of unstructured medical data: the impact of each preprocessing stage on classification

Abstract

Ähnliche Arbeiten

Autoren

Institutionen

Themen