why choose us

300×250 Ad Slot

Research Article: Large language models for breast cancer treatment planning: a blinded real-world evaluation of DeepSeek, ChatGPT, and oncologist recommendations

Date Published: 2026-06-30

Abstract:
Large language model (LLM) are increasingly explored for oncology decision support, yet their alignment with real-world clinical practice across varying disease complexities remains insufficiently characterized. This study aimed to evaluate and compare the accuracy, stability, and concordance of two advanced LLMs—DeepSeek V3.1 and ChatGPT-5—against experienced oncologists in generating breast cancer treatment plans within a specific clinical setting. This retrospective study compared the performance of DeepSeek V3.1 and ChatGPT-5 with senior oncologists using de-identified records from 213 breast cancer patients (Stages I–IV). To assess effectiveness, we implemented a multidimensional evaluation framework: accuracy was measured using a 5-point Likert scale by three independent, blinded expert reviewers; internal consistency was quantified via variance and coefficient of variation; and clinical concordance was evaluated using a structured five-level scoring system. Statistical analyses, including ANOVA and ordinal regression, were used to examine the impact of disease stage on AI-human agreement. Under standardized retrospective review conditions, LLM-generated recommendations demonstrated higher expert-rated guideline concordance and lower variability than historical real-world oncologist plans. Specifically, DeepSeek V3.1 achieved the highest expert-rated accuracy scores with minimal internal variance (4.91?±?0.36), outperforming both ChatGPT-5 (4.65?±?0.62) and clinicians (3.82?±?0.63, P <?0.001). While AI outputs exhibited high mutual consistency (74.2%), expert evaluations revealed a significant decline in AI-clinician agreement as disease stage advanced ( P <?0.001), particularly in Stage IV cases where clinicians prioritized real-world constraints such as financial toxicity. Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios. However, the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors. These findings support the role of LLMs as robust clinician-supervised decision-support tools while emphasizing the necessity of human judgment for individualized care.

Introduction:
Breast cancer is the most common malignancy in women globally and in China. In 2022, there were ?2.3 million new cases and 670,000 deaths worldwide, with ?357,200 new cases in China ( 1 , 2 ). Rising incidence highlights the urgent need for standardized, individualized treatment. Large language models (LLMs) have been studied and applied in the medical field, The integration of Artificial Intelligence, particularly Large Language Models (LLMs), is rapidly transforming healthcare demonstrating significant potential…

Read more

300×250 Ad Slot