Search papers, labs, and topics across Lattice.
4
0
9
0
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
Coding agents struggle to create complete and engaging games, with top performers barely reaching 41.46% success in end-to-end game generation.
LLMs struggle to accurately evaluate real human reasoning in mathematics, revealing a critical evaluation gap that challenges current assessment methods.
LLMs can learn new tasks without forgetting old ones, thanks to a memory-aware replay strategy that selectively rehearses important examples.