why choose us

300×250 Ad Slot

Research Article: Assessing retina-specific ophthalmic counseling generated by an early public large language model across different levels of clinical urgency

Date Published: 2026-07-01

Abstract:
To evaluate how the quality of retina-specific ophthalmology counseling provided by an early publicly available large language model (LLM) differs when advising patients with varying clinical characteristics and risk factors. Prospective, cross-sectional study. 18 ophthalmologists. Six patient vignettes were constructed with high- and low-urgency clinical scenarios for diabetic retinopathy (DR), retinal detachment (RD), and age-related macular degeneration (AMD). Based on these vignettes, an LLM (ChatGPT-3.5) was asked to provide written medical counseling in February 2024. Each response was rated on several metrics via 5-point Likert scale by 18 independent reviewers. Notably, readability was assessed both qualitatively via survey and quantitatively via Readable (an online readability tool that incorporates 5 different metrics). Counseling generated by the LLM was graded on accuracy, appropriate communication of urgency and empathy, readability, and potential for clinically significant harm. Counseling accuracy differed across levels of clinical urgency ( p =?0.002) but remained consistent between high- and low-urgency vignettes of AMD ( p =?0.081) and DR ( p =?0.5), albeit not for RD ( P <?0.001). Counseling urgency did not differ significantly from clinical urgency of all vignettes, except for the high-urgency AMD ( p =?0.013) and high-urgency RD ( p <?0.001). While counseling urgency did not significantly differ between high- and low-urgency vignettes of AMD ( p =?0.055) and RD ( p =?0.3), it did differ for the DR vignettes ( p <?0.001). Counseling empathy did not differ across clinical urgency ( p =?0.2). Four readability indices (e.g., Flesch Kincaid Grade Level) consistently indicated that college graduation would be required to understand every counseling output. Across all vignettes, the most common reasons for potential difficulty in understanding the counseling were having too much medical (49%) and non-medical (45%) terminology. The evaluated LLM-generated counseling outputs were largely similar across the sampled retinal vignettes with differing clinical urgency. Future studies should investigate the optimization of LLM prompting needed to garner counseling of consistent/appropriate accuracy, readability, empathy, and communication of urgency for specific conditions.

Introduction:
Large Language Model (LLM) chatbots have become an increasingly popular source of information due to their unprecedented ability to distill and present data. For example, since the release of ChatGPT in November 2022, the program has accrued over 300 million weekly active users worldwide ( 1 ). Given how often patients use the internet for health-related queries, it is imperative that the capabilities of LLMs in this context are carefully evaluated ( 2 ). Previous ophthalmology studies have demonstrated…

Read more

300×250 Ad Slot