Large Language Models in Patient Education: A Comparative Readability Analysis of Testicular Torsion Information Generated by ChatGPT-5 and Gemini 2.5 Pro
DOI:
https://doi.org/10.66288/actamedi.2026.92Keywords:
Artificial intelligence, Large language models, Testicular torsionAbstract
Background:Large language models (LLMs) are increasingly used to provide patients with health-related information. Although these systems have shown considerable potential in medical communication, the readability of AI-generated patient education materials remains an important concern, particularly for time-sensitive conditions such as testicular torsion, where delayed recognition and treatment may lead to irreversible testicular damage.
Objective:This study aimed to compare the readability of patient education materials on testicular torsion generated by ChatGPT-5 and Gemini 2.5 Pro using validated readability assessment tools.
Methods: A cross-sectional comparative study was conducted using standardized prompts submitted to ChatGPT-5 and Gemini 2.5 Pro. Forty patient information texts generated by each model (80 texts in total) were analyzed. Readability was evaluated using the Flesch Reading Ease Score (FRES), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index, Automated Readability Index (ARI), SMOG Index, Coleman–Liau Index, Dale–Chall Readability Score, Linsear Write Formula, and Spache Readability Formula. Statistical comparisons between groups were performed using independent samples t-tests or Mann–Whitney U tests, with statistical significance defined as p< 0.05.
Results: ChatGPT-5 generated significantly higher FRES values than Gemini 2.5 Pro (62.48 ± 3.84 vs. 45.44 ± 1.52, *p* < 0.001), indicating greater reading ease. Conversely, Gemini 2.5 Pro achieved significantly lower FKGL scores (6.46 ± 2.44 vs. 9.42 ± 1.48, *p* < 0.001), suggesting that its texts corresponded more closely to the sixth-grade reading level recommended for patient education. No statistically significant differences were observed for the ARI, Gunning Fog Index, Coleman–Liau Index, SMOG Index, Linsear Write Formula, Dale–Chall Readability Score, or Spache Readability Formula (all *p* > 0.05).
Conclusion: ChatGPT-5 and Gemini 2.5 Pro produced patient education materials with broadly comparable overall readability. ChatGPT-5 generated texts that were easier to read according to the FRES, whereas Gemini 2.5 Pro produced materials requiring a lower educational grade level according to the FKGL. Despite these differences, neither model consistently achieved the readability standards recommended for patient education. AI-generated educational materials should therefore undergo expert review before clinical use, particularly for emergency conditions such as testicular torsion.
References
1. Cummings JM, Boullier JA. The accuracy of physical examination in the diagnosis of testicular torsion. J Urol. 2002;168(5):2140-2.
2. Sharp VJ, Kieran K, Arlen AM. Testicular torsion: diagnosis, evaluation, and management. Am Fam Physician. 2013;88(12):835-40.
3. Ringdahl E, Teague L. Testicular torsion. Am Fam Physician. 2006;74(10):1739-43.
4. Berkman ND, Sheridan SL, Donahue KE, Halpern DJ, Crotty K. Low health literacy and health outcomes: an updated systematic review. Ann Intern Med. 2011;155(2):97-107. DOI: https://doi.org/10.7326/0003-4819-155-2-201107190-00005
5. Nielsen-Bohlman L, Panzer AM, Kindig DA, editors. Health Literacy: A Prescription to End Confusion. Washington (DC): National Academies Press; 2004. DOI: https://doi.org/10.17226/10883
6. Weiss BD. Health Literacy and Patient Safety: Help Patients Understand. 2nd ed. Chicago: American Medical Association Foundation; 2007.
7. Flesch R. A new readability yardstick. J Appl Psychol. 1948;32(3):221-33. DOI: https://doi.org/10.1037/h0057532
8. Kincaid JP, Fishburne RP Jr, Rogers RL, Chissom BS. Derivation of new readability formulas for Navy enlisted personnel. Res Branch Rep. 1975;8-75.
9. Gunning R. The Technique of Clear Writing. New York: McGraw-Hill; 1952.
10. Mc Laughlin GH. SMOG grading: A new readability formula. J Read. 1969;12(8):639-46.
11. Coleman M, Liau TL. A computer readability formula designed for machine scoring. J Appl Psychol. 1975;60(2):283-4. DOI: https://doi.org/10.1037/h0076540
12. Chall JS, Dale E. Readability Revisited: The New Dale-Chall Readability Formula. Cambridge: Brookline Books; 1995.
13. OpenAI. GPT-4 Technical Report. arXiv. 2023;2303.08774.
14. OpenAI. GPT-4.1 System Card. 2025.
15. Google DeepMind. Gemini 2.5: Advancing reasoning capabilities in multimodal AI. Google AI Blog. 2025.
16. Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel). 2023;11(6):887. DOI: https://doi.org/10.3390/healthcare11060887
17. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education. PLoS Digit Health. 2023;2(2):e0000198. DOI: https://doi.org/10.1371/journal.pdig.0000198
18. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388:1233-9. DOI: https://doi.org/10.1056/NEJMsr2214184
19. Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9:e45312. DOI: https://doi.org/10.2196/45312
20. Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-96. DOI: https://doi.org/10.1001/jamainternmed.2023.1838
21. Johnson D, Goodman R, Patrinely J, et al. Assessing the accuracy and reliability of AI-generated medical responses: a systematic review. NPJ Digit Med. 2024;7:45.
22. Rao A, Kim J, Kamineni M, et al. Evaluating ChatGPT as an adjunct for radiologic decision-making. Radiology. 2023;307(5):e230907. DOI: https://doi.org/10.1101/2023.02.02.23285399
23. Haver HL, Ambinder EB, Bahl M. Artificial intelligence in patient education: opportunities and challenges. Radiographics. 2024;44(1):e230145.
24. Lozano A, Martinez M, Perez J, et al. Readability of artificial intelligence-generated patient education materials across common medical conditions. BMC Med Inform Decis Mak. 2024;24:218.
25. D'Souza RS, Mathew RP, et al. Evaluation of readability and quality of ChatGPT-generated patient education materials. Cureus. 2024;16:e54012.
26. Eysenbach G. Improving the quality of Web-based health information. J Med Internet Res. 2002;4(1):e8. DOI: https://doi.org/10.2196/jmir.4.3.e17
27. National Institutes of Health. Clear Communication: An NIH Health Literacy Initiative. Bethesda (MD): NIH; 2023.
28. American Medical Association. Health Literacy and Patient Education Toolkit. Chicago: AMA; 2022.
29. World Health Organization. Health literacy development for the prevention and control of noncommunicable diseases. Geneva: World Health Organization; 2022.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Volkan Çelebi

This work is licensed under a Creative Commons Attribution 4.0 International License.