Search papers, labs, and topics across Lattice.
9
0
12
15
Output compression can slash inference costs by up to 3x, but input compression leads to higher costs and accuracy collapse鈥攁n unexpected trade-off for LLMs.
Shifting reasoning to the indexing stage can drastically reduce query latency while enhancing retrieval effectiveness through LLM-generated rationales.
Integrating semantic, acoustic, and engagement signals in music recommendations can boost performance by nearly 95%, challenging the status quo of opaque token-based systems.
Texture, not color, is the secret sauce behind fashion house identity, revealed by probing a multimodal CNN trained on decades of Vogue runway images.
LLMs can be made 20% more accurate by jointly attributing claims to sources and verifying them, rather than just verifying.
Multimodal agents can now plan more coherently and solve complex tasks thanks to a new anticipatory reasoning framework that forecasts short-horizon trajectories before acting.
Agentic RAG systems can be made significantly more efficient and accurate simply by adding a contextualization module and de-duplicating retrieved documents at test time.
A unified benchmark reveals the fragmented landscape of RAG security, highlighting vulnerabilities to knowledge-extraction attacks and paving the way for robust defense strategies.
LLM judges exhibit a surprising "blindness" to human-written summaries, increasingly preferring machine-generated content as the similarity to human references decreases, challenging their reliability in summarization tasks.