Search papers, labs, and topics across Lattice.
School of Computer Science and Engineering, Nanyang Technological University, Singapore
5
0
8
8
By intelligently offloading tasks to a large language model, PyroDash achieves a remarkable balance of accuracy and cost, slashing expenses from $49.36 to just $1.78 per example.
Evolving evaluation metrics can enhance LLM performance by up to 110% while maintaining transparency and robustness against manipulation.
A biased judge can silently disable skill retirement in self-evolving agents, leading to unnoticed performance degradation that can jeopardize deployment.
LLMs can teach themselves new tricks: a simple self-improvement loop, "Ratchet," lets a frozen LLM agent significantly boost its coding performance by autonomously managing its own library of natural-language skills.
End-to-end prompt optimization is often a waste of time and money, succeeding only when coaxing models into specific output formats they're already capable of.