Search papers, labs, and topics across Lattice.
5
0
8
0
Evolving evaluation metrics can enhance LLM performance by up to 110% while maintaining transparency and robustness against manipulation.
A biased judge can silently disable skill retirement in self-evolving agents, leading to unnoticed performance degradation that can jeopardize deployment.
Generated translation references outperform ground truth by 8.65 CEA100 points, showcasing a novel approach to literary translation challenges.
LLMs can teach themselves new tricks: a simple self-improvement loop, "Ratchet," lets a frozen LLM agent significantly boost its coding performance by autonomously managing its own library of natural-language skills.
Expert-written rules for coding agents are often useless or even harmful, with random constraints working just as well and negative constraints outperforming positive directives.