Search papers, labs, and topics across Lattice.
This paper introduces LLMRouter, a comprehensive infrastructure designed for the development, evaluation, and deployment of large language model (LLM) routers, addressing the inefficiencies of existing routing methods. By framing LLM routing as a sequential decision-making process with five key components, the authors create an automated pipeline for constructing routing supervision and evaluating routers based on both response quality and inference cost. Their empirical results demonstrate that learned routers significantly outperform fixed-model baselines, particularly under cost constraints, while also enhancing personalization through user-conditioned routing strategies.
Learned routers can outperform fixed-model baselines by 14.6%, revealing a new frontier in efficient LLM deployment.
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.