This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
875 posters, 25 topics, 3,440 authors, 1,061 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
March 25-28, 2026 | Tampa, FL, USA

P180
Artificial Intelligence
Title: Examining Large Language Models on Task of Improving Readability of Patient Education Materials for Hernia Surgery
Introduction -This study examines the replicability of creating simplified patient education materials (PEMs) with standardized prompts across three Large Language Models (LLM) platforms. PEMs are critical for informed consent and shared decision-making in surgery. Studies demonstrate that patient resources frequently exceed the sixth-grade reading level recommended by the AMA. LLMs may simplify existing PEMs, but their reliability, replicability, and risk of generating misinformation remain uncertain.
Methods and Procedures -Ventral Hernia PEM from SAGES was analyzed using a Flesch–Kincaid Grade Level (FKGL) and Flesch Reading Ease indices calculator. The text was then rewritten five times using a single prompt across three AI platforms: ChatGPT5 (LLM with open source access), NotebookLM (LLM restricted to user-uploaded sources), and OpenEvidence (Medical-oriented LLM). The standardized prompt directed each platform to output to a sixth-grade level utilizing the provided PEM PDF and no additional medical material. Rewritten texts were analyzed for FKGL and reviewers determined the rate of deviation from original PEM PDF information.
Results- SAGES Ventral hernia PEM was FKGL 8th grade. All three platforms improved readability. OpenEvidence had the best improvement in FKGL (mean 5.5, SD 0.5), followed by ChatGPT5 (mean 6, SD 0.8), then NotebookLM (mean 6.6, SD 0.6). Processing speed was fastest with NotebookLM (mean 19 seconds, SD 8sec) followed by OpenEvidence (mean 80 sec, SD 29sec), then ChatGPT5 (mean 69sec, SD 57sec). NotebookLM also had the lowest deviation from source material (mean 9%, SD 2%) compared to ChatGPT5 (mean 36%, SD 12%) and OpenEvidence (mean 51%, SD 10%).
Qualitative review revealed distinct behaviors. OpenEvidence introduced outside references despite instructions not to, increasing variability, but consistently produced the lowest FKGL. NotebookLM, although fastest and least prone to variation, oversimplified terminology (e.g., substituting “muscle” for “fascia”), but proved the most replicable by generating duplicate texts across multiple blind runs. ChatGPT5 produced outputs without citations and hallucinated postoperative timelines, presenting generally accepted, but unreferenced recommendations.
Conclusion-This study demonstrated the factors impacting potential use of LLMs for improving patient material readability. It highlights the potential strengths and weaknesses within LLM models that utilize different source material: ChatGPT5, OpenEvidence, and NotebookLM. Overall, NotebookLM had the fastest processing time and lowest variation rate likely due to its single-source model; while still meeting the 6th grade readability rating. As such, NotebookLM may offer clinicians and patients a more reliable and replicable option for improving access to patient information.
Reference:
SAGES. “Laparoscopic Ventral Hernia Repair Information from SAGES.” Society of American Gastrointestinal and Endoscopic Surgeons (SAGES), 19 Feb. 2022, www.sages.org/publications/patient-information/patient-information-for-laparoscopic-ventral-hernia-repair-from-sages/. Accessed 15 Sept. 2025.