Research Article: Benchmarking large language models on a Chinese radiation oncology technology examination-preparation question set: accuracy, consensus, and efficiency for AI assisted education
Abstract:
This study evaluated GPT-4o, GPT-5.4, and two DeepSeek platform configurations using 1, 053 examination-preparation questions from a publicly and commercially available 2025 exercise collection for the Chinese National Radiation Oncology Technology Qualification Examination (Intermediate Level). We assessed accuracy rates, identical-response coverage, accuracy consensus, and model-reported latency estimates.
Four configurations—GPT-4o, GPT-5.4, DS-Fast, and DS-Expert—were tested using an identical zero-shot Chinese prompt. Accuracy rates were summarized with Wilson 95% confidence intervals. Pairwise differences in the primary Chinese-language analysis were assessed using exact McNemar’s tests with Holm adjustment. A descriptive language-controlled analysis evaluated English translations of the same questions using an equivalent English prompt.
Overall accuracy rates were 56.7% (95% CI, 53.7–59.7) for GPT-4o, 66.1% (63.2–68.9) for GPT-5.4, 98.4% (97.4–99.0) for DS-Fast, and 99.8% (99.3–99.9) for DS-Expert. All six overall pairwise comparisons remained significant after Holm adjustment. DS-Fast and DS-Expert produced identical responses for 1, 034 of 1, 053 questions (98.2%), all of which were correct. GPT-5.4 and DS-Expert showed 100% conditional accuracy among shared responses, but their identical-response coverage was only 65.9% (694/1, 053). Model-reported latency estimates were lowest for DS-Fast (0.68 ± 1.00 seconds) and highest for DS-Expert (3.69 ± 2.24 seconds). Under the English-translated condition, overall accuracy rates were 62.6%, 61.5%, 66.5%, and 76.9%, respectively. DS-Expert remained the most accurate configuration, although its advantage over the GPT models was reduced.
DeepSeek configurations outperformed the GPT models on this Chinese-language examination-preparation question set, with DS-Expert achieving the highest accuracy. However, performance varied substantially by question language. These findings support further evaluation for answer verification and practice-question review, but they do not establish educational effectiveness.
Introduction:
In recent years, large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding, reasoning, and knowledge synthesis, attracting increasing attention regarding their potential applications in healthcare ( 1 – 3 ). Previous studies have evaluated LLM performance on standardized medical examinations, encompassing directions such as clinical decision support and medical education assistance ( 4 – 10 ). Multicenter studies have further identified the barriers and facilitators…
Read more