Research Article: Evaluating the accuracy and communication quality of large language models in Ewing sarcoma: a comparative analysis of ChatGPT, Claude, Gemini, DeepSeek, and Grok
Abstract:
Large language models (LLMs) are increasingly used to provide medical information, yet their performance in rare pediatric cancers remains largely unexplored. This study aimed to compare the clinical accuracy, comprehensiveness, and communication quality of five widely used LLMs in answering frequently asked questions about Ewing sarcoma.
Twelve representative questions covering diagnosis, treatment, prognosis, and psychosocial support were presented to ChatGPT (GPT-5.2), Claude Sonnet 4.5, Gemini 3, DeepSeek V3.2, and Grok 4. Two orthopedic oncology specialists independently evaluated each response using a 4-point Likert scale assessing clinical accuracy, completeness, clarity, and relevance. Qualitative assessments of empathy and communication quality were also performed. Statistical analyses included the Friedman, Wilcoxon signed-rank, Kruskal–Wallis, and Mann–Whitney U tests.
Significant differences were observed among the five LLMs ( p <?0.001). ChatGPT achieved the highest overall performance, followed by Claude and DeepSeek. DeepSeek demonstrated the greatest technical accuracy but lower communication quality, whereas ChatGPT provided the best balance between factual correctness and patient-friendly communication. Gemini and Grok produced more superficial responses with lower overall scores.
Current LLMs can support patient and family education in Ewing sarcoma but should not replace specialist consultation. Although ChatGPT and Claude demonstrated the most reliable overall performance, variability among models remains substantial. Further validation and disease-specific optimization are required before routine implementation in clinical practice.
Introduction:
Large language models (LLMs) are increasingly used to provide medical information, yet their performance in rare pediatric cancers remains largely unexplored. This study aimed to compare the clinical accuracy, comprehensiveness, and communication quality of five widely used LLMs in answering frequently asked questions about Ewing sarcoma.
Read more