Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Artificial Intelligence in Plastic Surgery Education: A Global Multimodel Benchmark of Large Language Models on the Plastic Surgery In-Service Training Examination

2026·0 Zitationen·Aesthetic Surgery Journal Open ForumOpen Access

Volltext beim Verlag öffnen

Zitationen

Autoren

2026

Jahr

Abstract

Background: Large language models (LLMs) are increasingly utilized in plastic surgery education. Previous studies have shown that flagship models can achieve high scores on medical examinations, including the Plastic Surgery In-Service Training Examination (PSITE). Yet evaluations often rely on single-shot accuracy of proprietary systems, neglecting stochastic variability and open-source or non-US alternatives. Objectives: The aim of this study was to comprehensively benchmark a globally representative cohort of 14 LLMs on the PSITE, assessing not only accuracy but also inter-run reliability and stochastic variability and to evaluate their role as educational tools in plastic surgery training. Methods: ) for reliability, and the coefficient of variation (CV) for stability. Stratified analyses assessed performance across clinical domains, proprietary vs open-source architectures, and paid vs free subscription tiers. Results: = 0.70). Stability varied, ranging from consistent error in Falcon H1 (CV = 0.00%) to erratic instability in Mistral Medium (Mistral AI, Paris, France) (CV = 32.2%). Conclusions: Contemporary LLMs possess substantial plastic surgery knowledge, yet meaningful disparities in reliability persist. Although proprietary models currently demonstrate superior reliability as educational tools, the presence of stochastic instability necessitates cautious adoption. Accuracy alone is insufficient to judge clinical utility; stability metrics are essential for selecting AI tools in surgical education.

Autoren

Institutionen

Themen

Artificial Intelligence in Healthcare and EducationDiversity and Career in MedicineSurgical Simulation and Training

Volltext beim Verlag öffnen

Artificial Intelligence in Plastic Surgery Education: A Global Multimodel Benchmark of Large Language Models on the Plastic Surgery In-Service Training Examination

Abstract

Ähnliche Arbeiten

Autoren

Institutionen

Themen