Search papers, labs, and topics across Lattice.
This paper introduces an innovative environment-free synthetic data generation method for training API-calling large language model (LLM) agents, addressing the scalability bottleneck caused by the need for high-quality trajectories from fully implemented environments. By utilizing LLMs as digital world models, the approach generates diverse tasks and simulates coherent API responses, resulting in a robust dataset for training. Evaluations on the AppWorld and OfficeBench benchmarks show that fine-tuning models on this synthetic data leads to significant performance improvements, validating the efficacy of LLM-based API simulation for diverse API ecosystems.
Training API-calling agents just got easier鈥攕ynthetic data generation using LLMs eliminates the need for complex environments while boosting performance.
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method generates trajectories mimicking interactions between an agent and a stateful environment. Specifically, an LLM first generates diverse tasks solvable with the provided APIs. A teacher agent then iteratively solves each task while an LLM simulator generates coherent synthetic API responses conditioned on the task context and simulation history. Finally, an LLM judge filters the trajectories to ensure the quality of the resulting dataset. We evaluate our approach on the challenging AppWorld and OfficeBench benchmarks, which include both information-retrieval and state-changing tasks. Fine-tuning models on our synthetic data yields significant performance gains, demonstrating that effective supervision for API-calling agents can be generated without any executable environment. Our results establish LLM-based API simulation as a practical, scalable solution for training agents across diverse API ecosystems.