Search papers, labs, and topics across Lattice.
2
0
3
1
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
LLMs' chain-of-thought reasoning often falls apart due to factual incompleteness, with errors compounding across multiple hops, as revealed by a new multi-hop QA dataset.