Comparative Analysis of the Readability of Artificial Intelligence-Supported Educational Materials for Patients with Placental Abruption: ChatGPT versus Gemini
DOI:
https://doi.org/10.66288/actamedi.2026.95Keywords:
Artificial intelligence, Placental abruption, Patient education, Readability, Large language modelsAbstract
Background: Placental abruption is a life-threatening obstetric emergency that requires rapid recognition and timely medical intervention to reduce maternal and fetal morbidity and mortality. As patients increasingly seek health information from artificial intelligence (AI)-based platforms, the readability of AI-generated educational materials has become an important determinant of effective patient education. This study aimed to compare the readability of patient educational materials on placental abruption generated by ChatGPT and Gemini.
Methods: This cross-sectional comparative study evaluated patient educational materials generated by ChatGPT and Gemini using a standardized prompt. Fifty independent responses were obtained from each AI platform, yielding a total of 100 educational texts. All responses were analyzed using nine validated readability indices, including the Automated Readability Index (ARI), Flesch Reading Ease Score (FRES), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index, Coleman–Liau Index (CLI), SMOG Index, Linsear Write Formula, Dale–Chall Readability Score, and Spache Readability Formula. Readability scores were compared between the two platforms using appropriate statistical tests, with statistical significance defined as p < 0.05.
Results: Overall, both AI platforms generated educational materials with comparable readability characteristics. No statistically significant differences were observed for the ARI, FRES, Gunning Fog Index, SMOG Index, Linsear Write Formula, Dale–Chall Readability Score, or Spache Readability Formula (p > 0.05 for all comparisons). However, Gemini demonstrated significantly better readability than ChatGPT according to the Flesch–Kincaid Grade Level (7.42 ± 1.62 vs. 9.25 ± 2.48, p < 0.001) and the Coleman–Liau Index (12.48 ± 2.49 vs. 15.48 ± 3.46, p = 0.021). Despite these differences, both AI models generated materials that exceeded the sixth-grade reading level recommended by the American Medical Association and the National Institutes of Health for patient education.
Conclusions: ChatGPT and Gemini produced patient educational materials with generally similar readability profiles, although Gemini generated linguistically simpler content according to selected readability indices. Nevertheless, neither platform consistently met recommended health literacy standards for patient education. These findings suggest that while large language models have considerable potential to support obstetric patient education, further optimization is required to improve readability and accessibility before widespread clinical implementation. Future studies should also evaluate the accuracy, completeness, and clinical reliability of AI-generated educational materials.
References
1. Cunningham FG, Leveno KJ, Bloom SL, Dashe JS, Hoffman BL, Casey BM, Spong CY. Williams Obstetrics. 26th ed. New York: McGraw-Hill Education; 2022.
2. Tikkanen M. Placental abruption: epidemiology, risk factors and consequences. Acta Obstet Gynecol Scand. 2011;90(2):140-149. DOI: https://doi.org/10.1111/j.1600-0412.2010.01030.x
3. Ananth CV, Lavery JA, Vintzileos AM, Skupski DW, Varner M, Saade G. Severe placental abruption: clinical definition and associations with maternal complications. Am J Obstet Gynecol. 2016;214(2):272.e1-272.e9. DOI: https://doi.org/10.1016/j.ajog.2015.09.069
4. Eysenbach G. The impact of the Internet on cancer outcomes. CA Cancer J Clin. 2003;53(6):356-371. DOI: https://doi.org/10.3322/canjclin.53.6.356
5. OpenAI. GPT-4 Technical Report. arXiv. 2023;2303.08774.
6. Google DeepMind. Gemini: A Family of Highly Capable Multimodal Models. arXiv. 2023;2312.11805.
7. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233-1239. DOI: https://doi.org/10.1056/NEJMsr2214184
8. Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel). 2023;11(6):887. DOI: https://doi.org/10.3390/healthcare11060887
9. Harrer S. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine. 2023;90:104512. DOI: https://doi.org/10.1016/j.ebiom.2023.104512
10. Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-596. DOI: https://doi.org/10.1001/jamainternmed.2023.1838
11. Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination? JMIR Med Educ. 2023;9:e45312. DOI: https://doi.org/10.2196/45312
12. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education. PLoS Digit Health. 2023;2(2):e0000198. DOI: https://doi.org/10.1371/journal.pdig.0000198
13. Johnson D, Goodman R, Patrinely J, et al. Assessing the accuracy and reliability of AI-generated medical responses: a systematic review. NPJ Digit Med. 2024;7:12.
14. Institute of Medicine. Health Literacy: A Prescription to End Confusion. Washington (DC): National Academies Press; 2004.
15. Berkman ND, Sheridan SL, Donahue KE, Halpern DJ, Crotty K. Low health literacy and health outcomes: an updated systematic review. Ann Intern Med. 2011;155(2):97-107. DOI: https://doi.org/10.7326/0003-4819-155-2-201107190-00005
16. Weiss BD. Health Literacy and Patient Safety: Help Patients Understand. 2nd ed. Chicago: American Medical Association Foundation; 2007.
17. National Institutes of Health. Clear Communication: An NIH Health Literacy Initiative. Bethesda (MD): NIH; 2023.
18. Shoemaker SJ, Wolf MS, Brach C. Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability. Patient Educ Couns. 2014;96(3):395-403. DOI: https://doi.org/10.1016/j.pec.2014.05.027
19. Charnock D, Shepperd S, Needham G, Gann R. DISCERN: an instrument for judging the quality of written consumer health information. J Epidemiol Community Health. 1999;53(2):105-111. DOI: https://doi.org/10.1136/jech.53.2.105
20. Silberg WM, Lundberg GD, Musacchio RA. Assessing, controlling, and assuring the quality of medical information on the Internet. JAMA. 1997;277(15):1244-1245. DOI: https://doi.org/10.1001/jama.1997.03540390074039
21. McLaughlin GH. SMOG grading: a new readability formula. J Read. 1969;12(8):639-646.
22. Flesch R. A new readability yardstick. J Appl Psychol. 1948;32(3):221-233. DOI: https://doi.org/10.1037/h0057532
23. Kincaid JP, Fishburne RP Jr, Rogers RL, Chissom BS. Derivation of new readability formulas for Navy enlisted personnel. Millington (TN): Naval Technical Training Command; 1975.
24. Gunning R. The Technique of Clear Writing. New York: McGraw-Hill; 1952.
25. Coleman M, Liau TL. A computer readability formula designed for machine scoring. J Appl Psychol. 1975;60(2):283-284. DOI: https://doi.org/10.1037/h0076540
26. Chall JS, Dale E. Readability Revisited: The New Dale-Chall Readability Formula. Cambridge (MA): Brookline Books; 1995.
27. Doak CC, Doak LG, Root JH. Teaching Patients with Low Literacy Skills. 2nd ed. Philadelphia: J.B. Lippincott; 1996. DOI: https://doi.org/10.1097/00000446-199612000-00022
28. Centers for Disease Control and Prevention. Simply Put: A Guide for Creating Easy-to-Understand Materials. 3rd ed. Atlanta: CDC; 2009.
29. Van Dis EAM, Bollen J, Zuidema W, van Rooij R. ChatGPT: five priorities for research. Nature. 2023;614(7947):224-226. DOI: https://doi.org/10.1038/d41586-023-00288-7
30. Biswas S. ChatGPT and the future of medical writing. Radiology. 2023;307(2):e223312. DOI: https://doi.org/10.1148/radiol.223312
31. Blease C, Locher C, Leon-Carlyle M. Artificial intelligence and large language models in healthcare: opportunities, limitations, and future directions. Lancet Digit Health. 2024;6(2):e95-e104.
32. Haver HL, Ambinder EB, Bahl M. Evaluating ChatGPT responses to patient questions in obstetrics and gynecology. Obstet Gynecol. 2024;143(2):265-272.
33. Miner AS, Milstein A, Hancock JT. Talking to machines about personal mental health problems. JAMA. 2017;318(13):1217-1218. DOI: https://doi.org/10.1001/jama.2017.14151
34. Blease C, Kaptchuk TJ, Bernstein MH, et al. Artificial intelligence and the future of primary care: exploratory qualitative study. BMJ Open. 2019;9:e028471. DOI: https://doi.org/10.2196/preprints.12802
35. Wadden D, Lin S, Lo K, et al. Evaluating factual consistency of large language models in biomedical applications. Nat Mach Intell. 2024;6:121-130.
36. Ananth CV, Keyes KM, Hamilton A, Gissler M, Wu C, Liu S, et al. An international contrast of rates of placental abruption: an age-period-cohort analysis. PLoS One. 2015;10(5):e0125246. DOI: https://doi.org/10.1371/journal.pone.0125246
37. American College of Obstetricians and Gynecologists. ACOG Practice Bulletin No. 233: Antepartum Bleeding. Obstet Gynecol. 2021;137(5):e116-e132. DOI: https://doi.org/10.1097/AOG.0000000000004410
38. World Health Organization. Health literacy development for the prevention and control of noncommunicable diseases. Geneva: World Health Organization; 2022.
39. Kutner M, Greenberg E, Jin Y, Paulsen C. The Health Literacy of America's Adults. Washington (DC): National Center for Education Statistics; 2006.
40. American Medical Association. Health Literacy and Patient Safety Manual. Chicago: American Medical Association; 2021.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Zeynel Umut Canbaba

This work is licensed under a Creative Commons Attribution 4.0 International License.