Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.
Critical assessment of transformer-based AI models for German clinical notes
28
Zitationen
12
Autoren
2022
Jahr
Abstract
Objective: , have recently received much attention. Currently, biomedical applications are primarily focused on the English language. While general-purpose German-language models such as GermanBERT and GottBERT have been published, adaptations for biomedical data are unavailable. This study evaluated the suitability of existing and novel transformer-based models for the German biomedical and clinical domain. Materials and Methods: We used 8 transformer-based models and pre-trained 3 new models on a newly generated biomedical corpus, and systematically compared them with each other. We annotated a new dataset of clinical notes and used it with 4 other corpora (BRONCO150, CLEF eHealth 2019 Task 1, GGPONC, and JSynCC) to perform named entity recognition (NER) and document classification tasks. Results: General-purpose language models can be used effectively for biomedical and clinical natural language processing (NLP) tasks, still, our newly trained BioGottBERT model outperformed GottBERT on both clinical NER tasks. However, training new biomedical models from scratch proved ineffective. Discussion: The domain-adaptation strategy's potential is currently limited due to a lack of pre-training data. Since general-purpose language models are only marginally inferior to domain-specific models, both options are suitable for developing German-language biomedical applications. Conclusion: General-purpose language models perform remarkably well on biomedical and clinical NLP tasks. If larger corpora become available in the future, domain-adapting these models may improve performances.
Ähnliche Arbeiten
"Why Should I Trust You?"
2016 · 14.750 Zit.
Coding Algorithms for Defining Comorbidities in ICD-9-CM and ICD-10 Administrative Data
2005 · 10.549 Zit.
A Comprehensive Survey on Graph Neural Networks
2020 · 8.957 Zit.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
2019 · 8.567 Zit.
High-performance medicine: the convergence of human and artificial intelligence
2018 · 8.083 Zit.
Autoren
Institutionen
- University of Bonn(DE)
- Fraunhofer Institute for Algorithms and Scientific Computing(DE)
- Bonn Aachen International Center for Information Technology(DE)
- Bielefeld University(DE)
- ZB MED - Information Centre for Life Sciences(DE)
- Humboldt-Universität zu Berlin(DE)
- Berlin Institute of Health at Charité - Universitätsmedizin Berlin(DE)
- Freie Universität Berlin(DE)
- Charité - Universitätsmedizin Berlin(DE)
- German Center for Lung Research(DE)