Research Article: Performance stability despite iteration: evaluating DeepSeek and ChatGPT on Chinese medical licensing examinations
Abstract:
Large Language Models (LLMs) hold substantial potential in medical education. In our previous work, we evaluated the performance of DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination (CNMLE). Following performance upgrades in December 2025, DeepSeek-V3.2 and ChatGPT-5.2 were released. This study aimed to longitudinally assess the performance evolution of DeepSeek (R1 vs. V3.2) and ChatGPT (4o vs. 5.2) using the 2024 CNMLE as a baseline and to explore their performance on the 2025 CNMLE.
We tested DeepSeek-V3.2 and ChatGPT-5.2 on 600 multiple-choice questions from the written part of the 2024 CNMLE, and compared the results with historical data from DeepSeek-R1 and ChatGPT-4o. The questions consisted of 4?units and were divided into low-difficulty and high-difficulty groups according to different difficulty levels. Additionally, the two latest LLMs were assessed on 600 questions from the written part of the 2025 CNMLE.
In the 2024 CNMLE, overall accuracy for the DeepSeek series (R1 vs. V3.2: 92.0% vs. 91.0%) and ChatGPT series (4o vs. 5.2: 87.2% vs. 89.3%) showed no significant differences (all p >?0.05). The accuracy gap between the two series narrowed from 4.8% (DeepSeek-R1 vs. ChatGPT-4o, p <?0.05) to 1.7% (DeepSeek-V3.2 vs. ChatGPT-5.2, p =?0.332). Subgroup analyses by unit and difficulty revealed no statistically significant differences, either in longitudinal comparisons between successive versions within each series or in cross-sectional comparisons between the latest versions of the two series. On the 2025 CNMLE, DeepSeek-V3.2 achieved significantly higher overall accuracy than ChatGPT-5.2 (94.7% vs. 89.3%, p <?0.05), with superior performance in Unit 2 and in both the low- and high-difficulty groups ( p <?0.05). Error analysis showed no significant differences across model versions in error type classification or clinical risk rating, despite a substantial proportion of high-risk clinical errors.
Benchmarked against the 2024 CNMLE, iterative updates did not yield significant performance gains for either LLM series. However, DeepSeek-V3.2 demonstrated a performance advantage on the 2025 CNMLE.
Introduction:
Large Language Models (LLMs) hold substantial potential in medical education. In our previous work, we evaluated the performance of DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination (CNMLE). Following performance upgrades in December 2025, DeepSeek-V3.2 and ChatGPT-5.2 were released. This study aimed to longitudinally assess the performance evolution of DeepSeek (R1 vs. V3.2) and ChatGPT (4o vs. 5.2) using the 2024 CNMLE as a baseline and to explore their performance on the 2025…
Read more