Search papers, labs, and topics across Lattice.
2
0
4
4
Evolving evaluation metrics can enhance LLM performance by up to 110% while maintaining transparency and robustness against manipulation.
A biased judge can silently disable skill retirement in self-evolving agents, leading to unnoticed performance degradation that can jeopardize deployment.