Search papers, labs, and topics across Lattice.
3
0
4
5
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
CMT-RAG achieves superior answer accuracy by transforming conversational memory into structured reasoning traces, enabling more effective multi-turn interactions.
LLMs' chain-of-thought reasoning often falls apart due to factual incompleteness, with errors compounding across multiple hops, as revealed by a new multi-hop QA dataset.