Araştırma Makalesi

ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5

Cilt: 9 24 Temmuz 2026
PDF İndir
TR EN

ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5

Öz

Objective: The aim of this study was to compare the accuracy, completeness, and readability of information related to whitening in dentistry provided by the large language models (LLM) ChatGPT-5, Gemini 2.5 Pro, DeepSeek v3.2, and Claude Sonnet 4.5. Methods: A total of 25 open-ended questions covering various aspects of dental whitening were prepared and presented to four different LLMs. The responses were recorded and evaluated independently by two restorative dental treatment specialists who were blinded to the source of the responses. Accuracy and completeness were evaluated using 5-point and 3-point Likert scales respectively. Readability was evaluated using the Flesch Reading Ease Score (FRES), the Flesch–Kincaid Grade Level (FKGL), and the Simple Measure of Gobbledygook (SMOG) index. In the statistical analyses, the Intraclass Correlation Coefficient (ICC), Shapiro-Wilk test, Kruskal-Wallis test, Dunn-Bonferroni test, and Spearman correlation analysis were used. Results: A statistically significant difference was identified among the artificial intelligence models regarding the accuracy of their responses to open-ended questions related to dental whitening (p<0.05). The accuracy score of DeepSeek v.3.2 (4.52±0.77) was determined to be statistically significntly higher than that of ChatGPT-5 (3.88±1.01). The FRES points for readability were found to be statistically significantly higher for DeepSeek v3.2 (42.02±7.63) compared to the mean points obtained for ChatGPT-5 (29.09±9.07), Gemini 2.5 Pro (33.92±9.24), and Claude Sonnet 4.5 (21.65±9.29). Conclusion: DeepSeek v3.2 exhibited better performance than ChatGPT-5 in terms of accuracy. No statistically significant differences were identified among the models regarding the completeness scores of responses to open-ended questions. However, the FRES results indicated that DeepSeek v3.2 generated texts with higher readability compared to the other chatbots. Overall, these findings suggest that DeepSeek v3.2 demonstrated superior performance in terms of both accuracy and readability.

Anahtar Kelimeler

Etik Beyan

This study was conducted in accordance with the ethical principles of the Helsinki Declaration and did not require Ethics Committee approval.

Kaynakça

  1. Thurzo A, Urbanova W, Novak B, Czako L, Siebert T, Stano P, et al. Where is artificial intelligence applied in dentistry? A systematic review and literature analysis. Healthcare. 2022;10:1234. doi:10.3390/healthcare10071269
  2. Acar AH. Can natural language processing serve as a consultant in oral surgery? J Stomatol Oral Maxillofac Surg. 2023;125:101724. doi:10.1016/j.jormas.2023.101724
  3. Dursun D, Bilici Geçer R. Can artificial intelligence models serve as patient information consultants in orthodontics? BMC Med Inform Decis Mak. 2024;24:211. doi:10.1186/s12911-024-02619-8
  4. Yu E, Chu X, Zhang W, et al. Large language models in medicine: applications, challenges, and future directions. Int J Med Sci. 2025;22:2792–2801. doi:10.7150/ijms.111780
  5. Xu X, Chen Y, Miao J. Opportunities, challenges, and future directions of large language models, including ChatGPT, in medical education: a systematic scoping review. J Educ Eval Health Prof. 2024;21:6. doi:10.3352/jeehp.2024.21.6
  6. Eggmann F, Blatz MB. ChatGPT: Chances and challenges for dentistry. Compend Contin Educ Dent. 2023;44(4):220–224.
  7. Taşyürek M, Adıgüzel Ö, Gündoğar M, Goncharuk-Khomyn M, Ortaç H. Comparative evaluation of the responses from ChatGPT-5, Gemini 2.5 Flash, and DeepSeek-V3.1 chatbots to patient inquiries about endodontic treatment in terms of accuracy, understandability, and readability. International Dental Research. 2025;15(3):91–95. doi:10.5577/intdentres.662.
  8. İlter Er Ö, Adıgüzel Ö, Taşyürek M, Ortaç H. Comparison of ChatGPT-5, Gemini 2.5 Pro, and DeepSeek-V3.1 chatbot responses to dental avulsion injuries. Health Sci Monit. 2025;3:e251201. doi:10.5577/hesmon.e251201

Ayrıntılar

Birincil Dil

İngilizce

Konular

Halk Sağlığı (Diğer)

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

24 Temmuz 2026

Gönderilme Tarihi

12 Mayıs 2026

Kabul Tarihi

22 Haziran 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 9

Kaynak Göster

APA
Cangül, S., Adıgüzel, Ö., Tunç, T., Taşyürek, M., & Ortaç, H. (2026). ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5. Acta Medica Nicomedia, 9, 1-12. https://doi.org/10.53446/actamednicomedia.1949463
AMA
1.Cangül S, Adıgüzel Ö, Tunç T, Taşyürek M, Ortaç H. ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5. Acta Med Nicomedia. 2026;9:1-12. doi:10.53446/actamednicomedia.1949463
Chicago
Cangül, Suzan, Özkan Adıgüzel, Tuba Tunç, Makbule Taşyürek, ve Hatice Ortaç. 2026. “ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5”. Acta Medica Nicomedia 9 (Temmuz): 1-12. https://doi.org/10.53446/actamednicomedia.1949463.
EndNote
Cangül S, Adıgüzel Ö, Tunç T, Taşyürek M, Ortaç H (01 Temmuz 2026) ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5. Acta Medica Nicomedia 9 1–12.
IEEE
[1]S. Cangül, Ö. Adıgüzel, T. Tunç, M. Taşyürek, ve H. Ortaç, “ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5”, Acta Med Nicomedia, c. 9, ss. 1–12, Tem. 2026, doi: 10.53446/actamednicomedia.1949463.
ISNAD
Cangül, Suzan - Adıgüzel, Özkan - Tunç, Tuba - Taşyürek, Makbule - Ortaç, Hatice. “ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5”. Acta Medica Nicomedia 9 (01 Temmuz 2026): 1-12. https://doi.org/10.53446/actamednicomedia.1949463.
JAMA
1.Cangül S, Adıgüzel Ö, Tunç T, Taşyürek M, Ortaç H. ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5. Acta Med Nicomedia. 2026;9:1–12.
MLA
Cangül, Suzan, vd. “ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5”. Acta Medica Nicomedia, c. 9, Temmuz 2026, ss. 1-12, doi:10.53446/actamednicomedia.1949463.
Vancouver
1.Suzan Cangül, Özkan Adıgüzel, Tuba Tunç, Makbule Taşyürek, Hatice Ortaç. ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5. Acta Med Nicomedia. 01 Temmuz 2026;9:1-12. doi:10.53446/actamednicomedia.1949463

images?q=tbn:ANd9GcSZGi2xIvqKAAwnJ5TSwN7g4cYXkrLAiHoAURHIjzbYqI5bffXt&s

"Acta Medica Nicomedia" Tıp dergisinde https://dergipark.org.tr/tr/pub/actamednicomedia adresinden yayımlanan makaleler açık erişime sahip olup Creative Commons Atıf-AynıLisanslaPaylaş 4.0 Uluslararası Lisansı (CC BY SA 4.0) ile lisanslanmıştır.