Research Article

Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study

Volume: 12 Number: 1 August 20, 2026
TR EN

Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study

Abstract

ABSTRACT Objective: Large language model chatbots are increasingly consulted for medical information. This study evaluated the accuracy and readability of chatbot responses to common patient questions on dry eye disease. Methods: This cross-sectional study analysed responses from four chatbots (ChatGPT 3.5, ChatGPT 4.0, Google Gemini, and Microsoft Copilot) to fifty standardised questions about dry eye disease. Two ophthalmologists independently rated accuracy on a five-point Likert scale, with inter-rater agreement measured by Cohen’s kappa. Readability was assessed using Flesch-Kincaid Grade Level, Gunning Fog Index, Coleman–Liau Index, Simple Measure of Gobbledygook, Flesch Reading Ease, and Reach score. Statistical tests included repeated-measures analysis of variance or Friedman tests with Bonferroni correction. Results: Agreement between raters was excellent (kappa = 0.88). Google Gemini showed the highest accuracy (4.84 ± 0.37), followed by Microsoft Copilot (4.76 ± 0.48), ChatGPT 4.0 (4.42 ± 0.50), and ChatGPT 3.5 (4.32 ± 0.59; p < 0.001). Gemini and Copilot significantly outperformed both ChatGPT versions. Readability differed significantly (p < 0.001). ChatGPT 4.0 produced the simplest texts, with the lowest grade levels, highest Flesch Reading Ease and broadest Reach. Gemini and ChatGPT 3.5 generated more complex responses, while Copilot showed intermediate values. Conclusions: Chatbots demonstrated complementary strengths. Gemini and Copilot were most accurate, whereas ChatGPT 4.0 was most readable and accessible. Chatbots may aid patient education in dry eye disease, but professional oversight remains necessary.

Keywords

References

  1. Pflugfelder SC, de Paiva CS. The Pathophysiology of Dry Eye Disease: What We Know and Future Directions for Research. Ophthalmol 2017;124(11S):S4-S13.
  2. Pflugfelder SC, De Paiva CS, Villarreal AL, Stern ME. Effects of sequential artificial tear and cyclosporine emulsion therapy on conjunctival goblet cell density and transforming growth factor- beta2 production. Cornea 2008; 27(1):64–9.
  3. Stapleton F, Alves M, Bunya VY, Jalbert I, Lekhanont K, Malet F, Na KS, Schaumberg D, Uchino M, Vehof J, Viso E, Vitale S, Jones L. TFOS DEWS II epidemiology report. Ocul surf 2017;15(3):334-365.
  4. Swoboda CM, Van Hulle JM, McAlearney AS, Huerta TR. Odds of talking to healthcare providers as the initial source of healthcare information: updated cross-sectional results from the Health Information National Trends Survey (HINTS). BMC Fam Pract. 2018;19(1):146.
  5. Rossettini G, Rodeghiero L, Corradi F, Cook C, Pillastrini P, Turolla A, Castellini G, Chiappinotto S, Gianola S, Palese A. Comparative accuracy of ChatGPT-4, Microsoft Copilot and Google Gemini in the Italian entrance test for healthcare sciences degrees: a cross-sectional study. BMC Med Educ. 2024;24(1):694.
  6. Aydın FO, Aksoy BK, Ceylan A, Akbaş YB, Ermiş S, Kepez Yıldız B, Yıldırım Y. Readability and Appropriateness of Responses Generated by ChatGPT 3.5, ChatGPT 4.0, Gemini, and Microsoft Copilot for FAQs in Refractive Surgery. Turk J Ophthalmol. 2024;54(6):313-317.
  7. Nichani PAH, Ong Tone S, AlShaker SM, Teichman JC, Chan CC. Use of Online Large Language Model Chatbots in Cornea Clinics. Cornea 2024;44(6):788-794.
  8. Owens OL, Leonard M. A Comparison of Prostate Cancer Screening Information Quality on Standard and Advanced Versions of ChatGPT, Google Gemini, and Microsoft Copilot: A Cross-Sectional Study. Am J Health Promot. 2025;39(5):766-776.

Details

Primary Language

English

Subjects

Ophthalmology

Journal Section

Research Article

Publication Date

August 20, 2026

Submission Date

January 8, 2026

Acceptance Date

May 12, 2026

Published in Issue

Year 2026 Volume: 12 Number: 1

APA
Kesimal, B., & Kocamış, S. İ. (2026). Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study. Akdeniz Tıp Dergisi, 12(1). https://doi.org/10.53394/akd.1859263
AMA
1.Kesimal B, Kocamış Sİ. Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study. Akd Med J. 2026;12(1). doi:10.53394/akd.1859263
Chicago
Kesimal, Bedia, and Sücattin İlker Kocamış. 2026. “Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study”. Akdeniz Tıp Dergisi 12 (1). https://doi.org/10.53394/akd.1859263.
EndNote
Kesimal B, Kocamış Sİ (August 1, 2026) Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study. Akdeniz Tıp Dergisi 12 1
IEEE
[1]B. Kesimal and S. İ. Kocamış, “Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study”, Akd Med J, vol. 12, no. 1, Aug. 2026, doi: 10.53394/akd.1859263.
ISNAD
Kesimal, Bedia - Kocamış, Sücattin İlker. “Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study”. Akdeniz Tıp Dergisi 12/1 (August 1, 2026). https://doi.org/10.53394/akd.1859263.
JAMA
1.Kesimal B, Kocamış Sİ. Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study. Akd Med J. 2026;12. doi:10.53394/akd.1859263.
MLA
Kesimal, Bedia, and Sücattin İlker Kocamış. “Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study”. Akdeniz Tıp Dergisi, vol. 12, no. 1, Aug. 2026, doi:10.53394/akd.1859263.
Vancouver
1.Bedia Kesimal, Sücattin İlker Kocamış. Appropriateness and Readability of Large Language Model Chatbot Responses to Frequently Asked Questions About Dry Eye Disease: Cross-Sectional Study. Akd Med J. 2026 Aug. 1;12(1). doi:10.53394/akd.1859263