Search papers, labs, and topics across Lattice.
This study investigates the transferability of prompt-side playbooks for tool-using language agents without retraining, employing a shared distill-validate-transfer protocol across multiple benchmarks. Results indicate that while playbook transfer can enhance performance under specific conditions, such as controlled greedy decoding, it is not universally effective, with only one of 135 route-level effects showing a significant advantage in a modest comparison. The findings highlight the necessity of target-side validation and the conditional nature of frozen transfer, suggesting it is not a default strategy for reuse.
Transferred playbooks can enhance agent performance, but their effectiveness hinges on specific conditions and requires careful validation.
Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On ALFWorld, transfer is beneficial under controlled greedy decoding and, in one near-budget-matched comparison, distilled guidance outperforms five fixed demonstrations. On TAU2-Bench, a prespecified aggregate contrast supports a modest average matched-domain advantage, but global Holm correction retains only one of 135 route-level effects; the remaining grid provides descriptive evidence of compatibility-sensitive heterogeneity. On XBench-DeepSearch, one artifact--runtime pairing preserves useful first-try heuristics while producing repeated queries, delayed stopping, and substantial cost inflation after a context-runtime shift. Across benchmarks, transferred and target-derived playbooks both require target-side validation of success, termination, protocol compatibility, and cost. Frozen transfer is therefore a conditional cold-start option, not a reuse-by-default strategy or a universally preferable alternative to target-side redistillation.