Search papers, labs, and topics across Lattice.
5
0
6
2
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
Assessment-free isolation can match the performance of full multi-agent assessment, revealing a surprising efficiency in weaker models.
DART boosts reasoning accuracy by up to 22.5 points while slashing thinking token usage by over 50%, all without requiring labeled training data.
ACOER reduces token generation by over 60% while boosting accuracy, solving the reward collapse problem that plagues traditional efficiency training methods.
LLMs' chain-of-thought reasoning often falls apart due to factual incompleteness, with errors compounding across multiple hops, as revealed by a new multi-hop QA dataset.