Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Comparison of GPT-5 and GPT-4o in Solving the Polish Centre for Medical Examinations (CEM) Gastroenterology Examination

2026·0 Zitationen·CureusOpen Access

Volltext beim Verlag öffnen

Zitationen

Autoren

2026

Jahr

Abstract

Both GPT-4o and GPT-5 exceeded the passing threshold for the CEM gastroenterology examination, demonstrating strong performance on a specialty-level medical assessment. Although overall accuracy was comparable, GPT-5 showed superior alignment between confidence and correctness, suggesting improved metacognitive reliability rather than a substantial gain in raw accuracy. These findings highlight the potential educational value of newer LLMs while underscoring important limitations, including the restricted sample size, exam-specific context, and lack of assessment of real-world clinical reasoning. Ethical considerations such as hallucinations, overconfidence, and inappropriate clinical reliance remain critical barriers to direct clinical deployment. Future research should focus on broader exam representativeness, task difficulty stratification, and controlled integration of LLMs into postgraduate medical education.

Autoren

Institutionen

Themen

Artificial Intelligence in Healthcare and EducationInnovations in Medical EducationClinical Reasoning and Diagnostic Skills

Volltext beim Verlag öffnen

Comparison of GPT-5 and GPT-4o in Solving the Polish Centre for Medical Examinations (CEM) Gastroenterology Examination

Abstract

Ähnliche Arbeiten

Autoren

Institutionen

Themen