Search papers, labs, and topics across Lattice.
This paper explores the transition of agentic systems based on large language models (LLMs) from research prototypes to real-world applications in fields like software engineering and finance. It highlights the challenges of robustness, safety, and reliability that arise during deployment, contrasting academic benchmarks with practical experiences. Through case studies in pharmaceutical discovery and financial systems, the authors identify successful design patterns and propose mitigation strategies for common failure modes, providing valuable insights for researchers and practitioners alike.
Real-world deployment of LLM-based agents reveals critical safety and reliability challenges that traditional benchmarks overlook.
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.