Intravitreal anti-VEGF therapy: Comparative evaluation of appropriateness and readability of large language model chatbots’ responses to frequently asked patient questions
Küçük Resim Yok
Tarih
2026
Yazarlar
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Elsevier Masson s.r.l.
Erişim Hakkı
info:eu-repo/semantics/closedAccess
Özet
Purpose To evaluate the appropriateness and readability of responses generated by large language model (LLM) chatbots to frequently asked patient questions regarding intravitreal anti-vascular endothelial growth factor (anti-VEGF) therapy. Methods Forty patient-centered anti-VEGF-related questions were developed by retinal specialists and posed in English to six LLM chatbots (ChatGPT-4.0, ChatGPT-5.2, Google Gemini 3, Microsoft Copilot, Grok 4, and Manus 1.6 Lite) under identical conditions. Responses were recorded verbatim and anonymized. Two ophthalmologists evaluated clinical appropriateness using a three-point Likert scale. Readability was assessed using five validated indices, and text length and time-based parameters were analyzed. Results None of the responses were classified as inappropriate. Gemini 3 demonstrated the highest rate of appropriate responses (97.5%), followed by ChatGPT-5.2 (90%), ChatGPT-4.0 (87.5%), and Manus 1.6 Lite (87.5%), while Copilot and Grok 4 showed lower appropriateness due to a higher proportion of partially appropriate responses (P = 0.033). Significant differences were observed across all readability indices (P < 0.001). Gemini 3 achieved the highest Flesch Reading Ease scores, indicating better patient accessibility, whereas Grok 4 produced more complex texts requiring higher educational levels. Manus 1.6 Lite generated the longest and most information-dense responses, while Gemini 3 demonstrated a more balanced profile between informational depth and readability. Conclusions While LLM chatbots generally provide clinically appropriate information on intravitreal anti-VEGF therapy, substantial model-dependent differences exist in readability and communication quality. LLMs should therefore be used as physician-supervised tools to support patient education rather than as standalone information sources. © 2026 Elsevier Masson SAS.
Açıklama
Anahtar Kelimeler
Large language models, Ophthalmology, Patient education, Readability, İntravitreal anti-VEGF
Kaynak
Journal Francais d'Ophtalmologie
WoS Q Değeri
Scopus Q Değeri
Q3
Cilt
49
Sayı
7












