206 posters, 13 topics, 5 sessions, 713 authors, 293 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
AAPS 104th Annual Meeting
May 2-5, 2026 | Lihue, HI

D85
LLM-assisted Screening For Plastic Surgery Systematic Reviews
Poster Presenter
Part of Topic
Education and Practice Management
LLM-Assisted Screening for Plastic Surgery Systematic Reviews
Neil Reddy, Nihal Sriramaneni, Nitya Devisetti, Justin Park, Pushti Shah, Christopher Didzbalis, Alex Wong
Purpose: Systematic reviews are essential in plastic surgery research, but abstract screening is labor-intensive and inconsistent. Large language models (LLMs) may accelerate screening, though optimal prompting strategies remain unclear. We evaluated LLM performance using PICO statements versus inclusion/exclusion (I/E) criteria and compared multiple models.
Methods: We analyzed 50 randomly sampled papers from three completed systematic reviews on plastic surgery and the lymphatic system. We tested abstract and full-text screening using PICO alone, I/E alone, or both combined. Papers were excluded in the combined condition only if both prompts rejected them. We measured sensitivity, specificity, and accuracy, comparing ChatGPT-5, SciSpace Agent, and Claude Sonnet 4.5.
Results: For abstract screening, mean sensitivity was 0.61±0.21 (PICO), 0.70±0.15 (I/E), and 0.85±0.14 (combined). Full-text sensitivity was 0.71±0.25 (PICO), 0.56±0.34 (I/E), and 0.90±0.08 (combined). Model sensitivities were 0.86 (ChatGPT), 0.79 (SciSpace), and 0.86 (Claude). Agreement with human screening was highest for ChatGPT (κ=0.71). Combining all three LLMs increased sensitivity to 0.93.
Conclusion: LLM-assisted screening can streamline systematic reviews but requires caution, as sensitivity fell below Cochrane's 0.99 threshold. Combined PICO and I/E prompting consistently improved sensitivity over single-prompt approaches. ChatGPT-5 and Claude Sonnet 4.5 performed similarly, with ChatGPT showing superior specificity. All models effectively filtered irrelevant articles but required human oversight to capture all eligible studies. Future research should optimize prompting strategies across diverse topics.
