Comparison of the Readability of Artificial Intelligence Applications in Periodontal Treatment

Authors

  • Halil İbrahim ŞIK HATAY MUSTAFA KEMAL UNİVERSİTY

DOI:

https://doi.org/10.66288/actamedi.2026.88

Keywords:

Artificial Intelligence, ChatGPT, Gemini, Readability, Periodontology

Abstract

Objective: This study aims to compare the readability of responses provided by two prominent Large Language Models (LLMs), ChatGPT and Gemini, to frequently asked questions by patients regarding periodontal diseases.

Materials and Methods: Ten common questions encountered by periodontology specialists at a university hospital were compiled. These questions were asked to both LLMs at 20 different times, and their responses were recorded. Readability was assessed using various indices, including the Automated Readability Index (ARI), Flesch Reading Ease, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, SMOG Index, Linsear Write Formula, and FORCAST. Statistical analysis was performed using independent samples t-tests, with p < 0.05 considered significant.

Results: Gemini produced texts with significantly higher reading levels—indicating greater complexity—compared to ChatGPT across several metrics: ARI (p=0.006), Flesch-Kincaid Grade Level (p=0.005), SMOG Index (p=0.003), and Linsear Write Grade Level (p=0.008). No statistically significant differences were found in Flesch Reading Ease (p=0.088), Gunning Fog Index (p=0.134), or Coleman-Liau Index (p=0.274).

Conclusion: The findings indicate that Gemini’s responses are generally more complex and challenging to read than those generated by ChatGPT. This suggests significant differences in how these models adapt their outputs for a general audience. Relying on a comprehensive set of readability formulas is essential when evaluating the suitability of AI-generated dental health content for patient education.

References

1. Pihlstrom BL, Michalowicz BS, Johnson NW. Periodontal diseases. Lancet (London, England). 2005;366(9499):1809-20. DOI: https://doi.org/10.1016/S0140-6736(05)67728-8

2. Kinane DF, Stathopoulou PG, Papapanou PN. Periodontal diseases. Nature reviews Disease primers. 2017;3:17038. DOI: https://doi.org/10.1038/nrdp.2017.38

3. Khodadady E, Mehrazmay RJSIJ. Evaluating two high intermediate EFL and ESL textbooks: a comparative study based on readability indices. 2017;1(3):93-102. DOI: https://doi.org/10.15406/sij.2017.01.00016

4. Gallagher T, Fazio X, Ciampa KJCJoERcdlé. A comparison of readability in science-based texts: Implications for elementary teachers. 2017;40(1):1-29.

5. Rakes TAJAE. A comparative readability study of materials used to teach adults to read. 1973;23(3):192-202. DOI: https://doi.org/10.1177/074171367302300302

6. Tussie C, Starosta AJBDJ. Comparing the dental knowledge of large language models. 2024:1-3. DOI: https://doi.org/10.21203/rs.3.rs-3974060/v1

7. Kuru HE, Aşık A, Demir DMJDT. Can artificial intelligence language models effectively address dental trauma questions? 2025. DOI: https://doi.org/10.1111/edt.13063

8. Biswas R, Mukhopadhyay A, Mukhopadhyay SJJoYms. Performance of large language models in fluoride-related dental knowledge: a comparative evaluation study of ChatGPT-4, Claude 3.5 Sonnet, Copilot, and Grok 3. 2025;42:53. DOI: https://doi.org/10.12701/jyms.2025.42.53

9. Taşyürek M, Adıgüzel Ö, Gündoğar M, et al. Comparative evaluation of the responses from ChatGPT-5, Gemini 2.5 Flash, and DeepSeek-V3. 1 chatbots to patient inquiries about endodontic treatment in terms of accuracy, understandability, and readability. 2025;15(Advanced Online). DOI: https://doi.org/10.5577/indentres.662

10. Giannakopoulos K, Kavadella A, Aaqel Salim A, et al. Evaluation of the performance of generative AI large language models ChatGPT, Google Bard, and Microsoft Bing Chat in supporting evidence-based dentistry: comparative mixed methods study. 2023;25:e51580. DOI: https://doi.org/10.2196/51580

11. Umer F, Batool I, Naved NJBo. Innovation and application of Large Language Models (LLMs) in dentistry–a scoping review. 2024;10(1):90. DOI: https://doi.org/10.1038/s41405-024-00277-6

12. Wu X, Cai G, Guo B, et al. A multi-dimensional performance evaluation of large language models in dental implantology: comparison of ChatGPT, DeepSeek, Grok, Gemini and Qwen across diverse clinical scenarios. 2025;25(1):1272. DOI: https://doi.org/10.1186/s12903-025-06619-6

13. Bedel HA, Bedel C, Selvi F, et al. Emergency Medicine Assistants in the Field of Toxicology, Comparison of ChatGPT-3.5 and GEMINI Artificial Intelligence Systems. Acta medica Lituanic. 2024;31(2):294-301. DOI: https://doi.org/10.15388/Amed.2024.31.2.18

14. Aydemir AM. Warfarin Use: A Readability Comparison of Gemini and ChatGPT: Warfarin Use: A Readability Comparison of Gemini and ChatGPT. Acta Medica Young Doctors. 2025;1(2):59-65.

15. Yavuz YF, Sevdimbaş AJAMYD. The Effectiveness of Large Language Models in the Diagnosis and Treatment of Glaucoma. 2025;1(1).

16. Şimşek SD, Şensöğüt CJAMYD. The Use of Readability Formulas in Polytrauma. 2025;1(1):2-9.

Downloads

Additional Files

Published

2026-07-08

How to Cite

ŞIK, H. İbrahim. (2026). Comparison of the Readability of Artificial Intelligence Applications in Periodontal Treatment. Acta Medica Young Doctors, 2(4). https://doi.org/10.66288/actamedi.2026.88

Similar Articles

1 2 > >> 

You may also start an advanced similarity search for this article.