why choose us

300×250 Ad Slot

Research Article: Benchmark evaluation of multi-modal large language models for ophthalmic diagnosis in real world

Date Published: 2026-06-22

Abstract:
Multimodal large language models (MLLMs) are increasingly demonstrating substantial potential in the medical domain, particularly in image-intensive specialties such as ophthalmology. Although cutting-edge models like ChatGPT-4o and Qwen-VL 2.5 have shown strong performance on general-domain tasks, real-world clinical benchmarks for rigorously assessing their diagnostic capabilities in specialized medical contexts remain limited. To address this gap, we constructed a carefully curated benchmark dataset comprising 295 pathologically confirmed ophthalmic cases with representative clinical presentations. Using this dataset, we systematically evaluated nine leading MLLMs, including both open-source and proprietary models. The results showed that models such as HAIBU-ReMUD and ChatGPT-4o achieved comparatively strong diagnostic accuracy and consistency, with performance in some settings approaching that of human experts. These findings suggest that current MLLMs are showing encouraging feasibility for real-world clinical applications and provide a basis for further investigation of their integration into ophthalmology practice.

Introduction:
Multi-modal large language models (MLLMs) are increasingly recognized for their potential to revolutionize clinical interpretation of medical images, offering advanced capabilities in diagnostic classification, lesion localization, and cross-modal reasoning ( 1 – 5 ). Recent advancements in MLLMs have enabled these systems to simultaneously process both textual and visual inputs, making them particularly well suited for tasks that require integrating narrative clinical descriptions with high-resolution medical…

Read more

300×250 Ad Slot