Search papers, labs, and topics across Lattice.
This study introduces GatorOnco, an agentic generative large language model specifically designed for colorectal cancer treatment planning, which integrates 282 billion tokens of biomedical text and employs a novel domain-adaptation method. By utilizing a retrieval-augmented generation approach, GatorOnco dynamically incorporates time-sensitive clinical guidelines, significantly outperforming open-source LLMs and achieving expert-level performance comparable to oncologists in a clinical evaluation. The model not only excelled in readability and completeness but also maintained correctness and safety on par with expert oncologists, highlighting its potential for high-stakes clinical applications.
GatorOnco outshines traditional models by achieving expert-level treatment planning performance while enhancing readability and completeness in colorectal cancer care.
Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnostic tasks, their adoption for high-stakes treatment planning is hindered by complex reasoning, adherence to timely clinical guidelines, and safety concerns. In this study, we present GatorOnco, an agentic LLM for colorectal cancer (CRC) treatment planning. GatorOnco is developed using a total of 282 billion tokens of biomedical text, including healthcare system-scale clinical text comprising 166 billion tokens from UF Health. We implemented a domain-adaptation method that integrates pre-training, model merging, a two-stage post-training approach, and agent-based reinforcement learning. An agentic retrieval-augmented generation (RAG) approach dynamically integrates time-sensitive clinical guidelines into the reasoning process. In a blind, randomized clinical evaluation conducted by five UF Health oncologists, GatorOnco significantly outperformed open-source LLMs (P < 0.01) and achieved expert-level performance comparable to UF Health oncologists. Compared with expert oncologists, GatorOnco received significantly higher ratings for readability (4.46 vs. 4.19, P < 0.01) and completeness (3.91 vs. 3.52, P < 0.01), while showing statistically comparable performance in correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999). These findings demonstrate that integrating agentic reasoning with large-scale domain adaptation can help bridge the gap for generative AI in high-stakes cancer treatment planning.