Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.
Scalable information extraction from free text electronic health records using large language models
23
Zitationen
7
Autoren
2025
Jahr
Abstract
BACKGROUND: A vast amount of potentially useful information such as description of patient symptoms, family, and social history is recorded as free-text notes in electronic health records (EHRs) but is difficult to reliably extract at scale, limiting their utility in research. This study aims to assess whether an "out of the box" implementation of open-source large language models (LLMs) without any fine-tuning can accurately extract social determinants of health (SDoH) data from free-text clinical notes. METHODS: We conducted a cross-sectional study using EHR data from the Mass General Brigham (MGB) system, analyzing free-text notes for SDoH information. We selected a random sample of 200 patients and manually labeled nine SDoH aspects. Eight advanced open-source LLMs were evaluated against a baseline pattern-matching model. Two human reviewers provided the manual labels, achieving 93% inter-annotator agreement. LLM performance was assessed using accuracy metrics for overall, mentioned, and non-mentioned SDoH, and macro F1 scores. RESULTS: . openchat_3.5 was the best-performing model, surpassing the baseline in overall accuracy across all nine SDoH aspects. The refined pipeline with prompt engineering reduced hallucinations and improved accuracy. CONCLUSIONS: Open-source LLMs are effective and scalable tools for extracting SDoH from unstructured EHRs, surpassing traditional pattern-matching methods. Further refinement and domain-specific training could enhance their utility in clinical research and predictive analytics, improving healthcare outcomes and addressing health disparities.
Ähnliche Arbeiten
A SIMPLE METHOD OF ESTIMATING FIFTY PER CENT ENDPOINTS12
1938 · 20.096 Zit.
Health Behavior and Health Education: Theory, Research, and Practice
1992 · 13.311 Zit.
Self-Rated Health and Mortality: A Review of Twenty-Seven Community Studies
1997 · 8.571 Zit.
Human gut microbiome viewed across age and geography
2012 · 7.795 Zit.
Social Relationships and Health
1988 · 7.020 Zit.