This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
119 posters, 6 topics, 524 authors, 243 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
29-30 June, 2026 | QEII Centre, Westminster

207
Hania Paverd, Zeyu Gao, Golnar Mahani, Margarete Fabre, Sarah Burge, Matthew Hoare, Mireia Crispin-Ortuzar
Early Cancer Institute, Department of Oncology, University of Cambridge, Centre for Genomics Research, Discovery Sciences, BioPharmaceuticals R&D, AstraZeneca, Cancer Research UK Cambridge Centre, University of Cambridge
AI Education and research: examples of proof of concept or AI in development, technical advances, teaching approaches or pre-clinical testing
Vision and Goals:
- Electronic health records contain rich clinical data but are often stored as unstructured text, making manual curation labor-intensive and error-prone.
- Large language models (LLMs) can automate extraction of structured data from clinical reports, yet most studies focus on a single modality, use proprietary models, and overlook challenges like output formatting.
- This study benchmarks open-source LLMs and constrained decoding methods across multiple report types to reconstruct disease trajectories in liver transplant patients.
Does constrained decoding help?
- Key finding: Constrained decoding leads to perfect output format adherence.
How do LLMs perform?
- Key finding: Llama 70B outperforms regex, OpenBioLLM and Llama 8B.
What can we learn from LLM-extracted data?
- Key finding: LLMs can identify clinically-validated risk factors of liver cancer.
- Key finding: LLM-extracted data can serve to reconstruct patient timelines from multiple report modalities.
Key message: Open-source LLMs with constrained decoding enable robust, scalable, automated data extraction from complex, multi-source clinical records.