Search papers, labs, and topics across Lattice.
This study systematically evaluates programmatic tool calling (PTC) against native JSON tool calling across 14 language models on the BFCL v4 benchmark, revealing that PTC often surpasses JSON in performance. The GPT-5.6 family, in particular, shows a notable 10.6% improvement over the JSON baseline, highlighting the effectiveness of exposing tools as typed Python stubs for direct invocation. Additionally, PTC demonstrates resilience under context rot conditions, maintaining performance where JSON tool calling degrades.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across current and prior model generations under real-world task conditions has not been conducted. In this work, we empirically compare programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4. In the programmatic tool calling paradigm, tools are exposed as typed Python stubs that the model invokes through code, with execution and results handled in a single agent turn. Programmatic tool calling matches or exceeds native JSON tool calling in 11 of 14 models on BFCL v4, with the GPT-5.6 family achieving a 10.6% improvement over the JSON tool calling baseline. Further, it matches or outperforms baseline in 13 of 14 models under parallel fan-out, and holds stable under context rot conditions where baseline degrades 2.3% on average. Our results demonstrate that programmatic tool calling is a viable and robust alternative to JSON tool calling, with performance tracking model capability across release generations.