why choose us

300×250 Ad Slot

Research Article: Freely accessible large language models for parent education in pediatric immune thrombocytopenia: an expert-rated cross-sectional study of safety, readability, and guideline concordance

Date Published: 2026-07-01

Abstract:
Parents of children with immune thrombocytopenia (ITP) increasingly turn to freely available large language models (LLMs) when they need quick explanations. In pediatric hematology, such advice must be clear, safe, and aligned with clinical guidance, not simply fluent. To assess the quality, safety, readability, and guideline/reference concordance of freely accessible LLM responses to parent-oriented questions about childhood ITP, as an evaluation of parent-facing educational content rather than diagnostic or predictive AI performance. Forty standardized parent-facing ITP questions were submitted once to GPT-5.3-mini, Gemini 3 Flash, and Claude Sonnet 4.6 through free public web interfaces. Reporting followed STROBE and was cross-mapped to relevant TRIPOD-LLM items. Three model-name-blinded clinical reviewers independently rated 120 responses using the study-specific, expert-developed seven-domain AI-ITP Parent Response Score (AI-ITP-PRS; range, 7–35) and recorded unsafe-content and hallucination flags. Readability was measured with automated indices. Paired model comparisons used the Friedman test with Holm-adjusted Wilcoxon analyses; repeated-measures ANOVA and mixed-effects models were used as sensitivity analyses. All 360 reviewer-level assessments were complete. The unweighted clinical-communication composite AI-ITP-PRS differed significantly across models (Friedman chi-square?=?63.52, p <?0.001; Kendall's W?=?0.794). Gemini 3 Flash had the highest mean composite score (32.78?±?2.54), followed by Claude Sonnet 4.6 (30.60?±?1.39) and GPT-5.3-mini (29.98?±?1.52), but this difference was driven primarily by completeness, comprehensibility, empathy/supportiveness, and objective readability. Medical accuracy, guideline/reference concordance, safety/emergency triage, and low harmful misinformation risk were largely similar across models. Two response-level unsafe-content and hallucination events were detected, both in Gemini 3 Flash responses (2/40, 5.0%; exact 95% CI, 0.6%–16.9%). The matched categorical comparison was not statistically significant, so these rare-event findings were interpreted descriptively and as hypothesis-generating. The events involved overgeneralized autoimmune-testing advice and unsupported MMR vaccination counseling. Freely accessible LLMs can produce supportive, high-scoring parent-facing explanations about childhood ITP. However, the highest composite score in this single-response snapshot reflected stronger communication performance rather than definitive superiority across safety-sensitive domains. This study detected high-impact safety failure modes but cannot precisely estimate true safety-event rates, establish robust rare-event safety differences between models, or define stable model behavior across repeated generations.

Introduction:
Parents of children with immune thrombocytopenia (ITP) increasingly turn to freely available large language models (LLMs) when they need quick explanations. In pediatric hematology, such advice must be clear, safe, and aligned with clinical guidance, not simply fluent.

Read more

300×250 Ad Slot