Search papers, labs, and topics across Lattice.
Cohere Labs, University of Amsterdam, Zeta Alpha Vector
8
72
10
23
Smaller models can outperform larger counterparts in linguistic reasoning tasks, challenging assumptions about model size and capability.
Confidence calibration in reasoning models can be dramatically improved by aligning supervision with the model's state, cutting ECE by over half on challenging benchmarks.
Static exposure scores miss the mark on the dynamic realities of AI's impact on work, risking misguided policy decisions.
Decomposing complex tasks into verifiable checklists unlocks more effective reinforcement learning, but only if you can avoid the pitfalls of reward hacking and verifier bias.
Forget brute-force scaling: Tiny Aya proves a 3B parameter model can achieve state-of-the-art multilingual performance with clever training and region-aware specialization.
Generic reasoning hurts machine translation, but a new structured reasoning approach鈥攚ith iterative drafting, refinement, and revision鈥攗nlocks significant gains.
Chatbot Arena, the go-to LLM leaderboard, is systematically gamed by undisclosed private testing and data access advantages, leading to biased rankings and overfitting.
Command A shows how to build an enterprise-grade LLM that balances performance, efficiency, and multilingual capabilities using decentralized training and model merging.