Araştırma Makalesi

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

Cilt: 7 Sayı: 3 16 Eylül 2026
PDF İndir
TR EN

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

Öz

Background: Although artificial intelligence (AI) has been used in patient education for some time, the accuracy, reliability, and clinical appropriateness of AI-generated medical content remain inadequately defined and continue to be debated. This study aimed to evaluate the reliability, readability, and comprehensibility of responses generated by ChatGPT 5.2 to the most frequently asked questions related to sleeve gastrectomy, and to assess the responses by surgeons with varying levels of clinical expertise. Methods: Twenty-four questions regarding sleeve gastrectomy were asked to ChatGPT 5.2, and responses were evaluated using a Likert-Scale by three general surgeons with varying levels of expertise. The readability and comprehensibility analysis was also conducted, using Flesch–Kincaid Grade Level and Flesch Reading Ease Scores. Results: Although excellent intra-rater reliability was observed, inter-rater reliability was poor. The evaluator with the greatest clinical expertise assigned the highest mean score (2.92±0.776), whereas the evaluator possessing predominantly theoretical knowledge assigned the lowest mean score (2.58±0.881). In the readability analysis, responses in all subcategories were classified as “difficult to read,” whereas only the responses within the postoperative course category were categorized as “fairly difficult”. Conclusions: ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable. Nevertheless, the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability observed among evaluators with different levels of clinical experience underscore the indispensable role of clinical expertise and specialized professional guidance in surgical practice.

Anahtar Kelimeler

ChatGPT, bariatric surgery, sleeve gastrectomy, frequently asked questions, artificial intelligence, patient education.

Destekleyen Kurum

The authors received no financial support for the research, authorship, and/or publication of this article.

Etik Beyan

The study did not involve human participants, patient records, or identifiable personal data. Consequently, in accordance with the findings of analogous studies documented in the extant literature, the institutional review board was deemed unnecessary, as no issues pertaining to patient privacy or personal data protection were involved

Teşekkür

None.

Kaynakça

  1. Velardi AM, Anoldo P, Nigro S, Navarra G. Advancements in Bariatric Surgery: A Comparative Review of Laparoscopic and Robotic Techniques. J Pers Med. 2024;14(2):151.
  2. Xia Q, Campbell JA, Ahmad H, Si L, de Graaff B, Palmer AJ. Bariatric surgery is a cost-saving treatment for obesity-A comprehensive meta-analysis and updated systematic review of health economic evaluations of bariatric surgery. Obes Rev. 2020;21(1):e12932.
  3. Rajan R, Sam-Aan M, Kosai NR, Shuhaili MA, Chee TS, Venkateswaran A, et al. Early outcome of bariatric surgery for the treatment of type 2 diabetes mellitus in super-obese Malaysian population. J Minim Access Surg. 2020;16(1):47-53.
  4. Widjaja J, Pan H, Dolo PR, Yao L, Li C, Shao Y, et al. Short-Term Diabetes Remission Outcomes in Patients with BMI ≤ 30 kg/m2 Following Sleeve Gastrectomy. Obes Surg. 2020;30(1):18-22.
  5. Erdem H, Sisik A. The Reliability of Bariatric Surgery Videos in YouTube Platform. Obes Surg. 2018;28(3):712-6.
  6. Lünse S, Wisotzky EL, Höhn J, Paasch C, Meyer F, Hunger R, et al. ChatGPT in general surgery: a cross-sectional study assessing its response to patient questions. Ann Med Surg (Lond). 2025;87(11):7088-94.
  7. Munir MM, Endo Y, Ejaz A, Dillhoff M, Cloyd JM, Pawlik TM. Online artificial intelligence platforms and their applicability to gastrointestinal surgical operations. J Gastrointest Surg. 2024;28(1):64-9.
  8. Horesh N, Emile SH, Gupta S, Garoufalia Z, Gefen R, Zhou P, et al. Comparing the Management Recommendations of Large Language Model and Colorectal Cancer Multidisciplinary Team: A Pilot Study. Dis Colon Rectum. 2025;68(1):41-7.
  9. Maron CM, Emile SH, Horesh N, Freund MR, Pellino G, Wexner SD. Comparing answers of ChatGPT and Google Gemini to common questions on benign anal conditions. Tech Coloproctol. 2025;29(1):57.
  10. Sharma S, Pajai S, Prasad R, Wanjari MB, Munjewar PK, Sharma R, et al. A Critical Review of ChatGPT as a Potential Substitute for Diabetes Educators. Cureus. 2023;15(5):e38380.

Kaynak Göster

APA
Türkoğlu, F., Gencer, E. N., & Erdoğan, E. (2026). Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2. Archives of Current Medical Research, 7(3), 737-746. https://doi.org/10.47482/acmr.1958370
AMA
1.Türkoğlu F, Gencer EN, Erdoğan E. Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2. Arch Curr Med Res. 2026;7(3):737-746. doi:10.47482/acmr.1958370
Chicago
Türkoğlu, Furkan, Elif Nur Gencer, ve Emre Erdoğan. 2026. “Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2”. Archives of Current Medical Research 7 (3): 737-46. https://doi.org/10.47482/acmr.1958370.
EndNote
Türkoğlu F, Gencer EN, Erdoğan E (01 Eylül 2026) Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2. Archives of Current Medical Research 7 3 737–746.
IEEE
[1]F. Türkoğlu, E. N. Gencer, ve E. Erdoğan, “Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2”, Arch Curr Med Res, c. 7, sy 3, ss. 737–746, Eyl. 2026, doi: 10.47482/acmr.1958370.
ISNAD
Türkoğlu, Furkan - Gencer, Elif Nur - Erdoğan, Emre. “Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2”. Archives of Current Medical Research 7/3 (01 Eylül 2026): 737-746. https://doi.org/10.47482/acmr.1958370.
JAMA
1.Türkoğlu F, Gencer EN, Erdoğan E. Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2. Arch Curr Med Res. 2026;7:737–746.
MLA
Türkoğlu, Furkan, vd. “Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2”. Archives of Current Medical Research, c. 7, sy 3, Eylül 2026, ss. 737-46, doi:10.47482/acmr.1958370.
Vancouver
1.Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan. Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2. Arch Curr Med Res. 01 Eylül 2026;7(3):737-46. doi:10.47482/acmr.1958370