Search papers, labs, and topics across Lattice.
3
7
6
4
ToFu achieves superior token efficiency and cost-effectiveness in agentic coding, all while empowering researchers with a transparent, modifiable framework.
Stop LLMs from drifting to English when reasoning in other languages: language-adaptive RL can guide them to stay consistent without sacrificing performance.
By recognizing that not all tokens are created equal, D2PO offers a simple temporal weighting fix that boosts DPO alignment scores by up to 9.7 points.