Search papers, labs, and topics across Lattice.
This paper conducts a scoping review of the current landscape of agentic AI in medicine, analyzing 557 studies that focus on goal-directed task execution, tool use, and multi-agent collaboration within clinical contexts. The findings reveal that while there is significant progress in applying AI to complex medical tasks, existing evaluation practices are inadequate for clinical translation, often relying on public benchmarks and simulated environments. The authors highlight the need for improved definitions and evaluation methods to ensure the reliability and safety of AI systems in real-world medical applications.
Current evaluation practices for agentic AI in medicine are misaligned with clinical needs, risking the reliability of these systems in real-world applications.
Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.