Search papers, labs, and topics across Lattice.
6
0
7
As human oversight wanes, LRMs could autonomously evolve, but this shift introduces significant risks like reward hacking and feedback drift.
EvolveNet reveals that decentralized evolution of agent harnesses can lead to substantial performance gains by leveraging localized experience rather than relying on centralized optimization.
A single misleading document can drastically reduce deep research agents' accuracy by up to 88%, exposing a critical vulnerability in their evidential reasoning capabilities.
Language corrections in PhysClaw-0 not only enhance robot autonomy but also boost success rates by over 35% while slashing human oversight time.
Harnesses can evolve in real-time during evaluation, leading to significant performance gains without retraining the underlying model.
Forget prompt engineering: MOSS lets autonomous agents rewrite their own source code to fix bugs and improve performance in production.