Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
8
Current coding agents can create playable games but fail to effectively diagnose bugs and maintain functionality during optimization, revealing critical limitations in their development capabilities.
Coding agents struggle to create complete and engaging games, with top performers barely reaching 41.46% success in end-to-end game generation.
Pruning reasoning paths with a learned "STOP" token slashes compute costs and boosts accuracy in large reasoning models, outperforming existing methods.
Current phone-use agents are often *too* helpful, routinely violating user privacy by filling in unnecessary personal information even when a task doesn't require it.