why choose us

300×250 Ad Slot

Research Article: Knowledge localization is associated with higher performance of domestic large language models in a Chinese radiation oncology examination

Date Published: 2026-06-17

Abstract:
The rapid evolution of large language models has demonstrated human-level proficiency in general English-language medical examinations, yet their applicability in highly specialized, non-English clinical domains remains insufficiently explored. To bridge this gap, this study benchmarked leading domestic (Qwen 3 Max, DeepSeek V3.2) and international (GPT-5, Gemini 2.5 Flash) models against a board-certified Chinese radiation oncologist, using the National Intermediate Professional Title Examination for Radiation Oncology in China. Statistical comparisons were conducted using mixed-effects logistic regression and McNemar's test for paired data. Domestic models demonstrated strong performance, with Qwen achieving an accuracy of 86.30%, which was higher than that of the single physician reference participant (adjusted P = 0.020). However, because only one human participant was included, this comparison should be interpreted as an illustrative reference rather than evidence of general superiority over radiation oncologists. In contrast, international models exhibited a marked performance decline in this specific context. Detailed cognitive stratification revealed a distinct reasoning-recall dissociation in international models, which maintained relatively robust clinical reasoning capabilities but failed significantly in localized knowledge retrieval. Notably, translating the examination into English did not bridge this performance gap for international models, while revealing a significant language penalty for certain domestic architectures (e.g., DeepSeek, P = 0.013). Granular error analysis confirmed that the majority of failures in international models like GPT-5 stemmed from discrepancies between Western and Chinese clinical guidelines rather than intrinsic reasoning deficits or hallucinations. These findings suggest that, within this examination-based Chinese radiation oncology benchmark, alignment with regional clinical standards may be a major contributor to model performance. However, differences in model architecture, training data, post-training, interface configuration, and potential test-set contamination may also contribute.

Introduction:
The rapid evolution of large language models has demonstrated human-level proficiency in general English-language medical examinations, yet their applicability in highly specialized, non-English clinical domains remains insufficiently explored.

Read more

300×250 Ad Slot