Search papers, labs, and topics across Lattice.
5
0
6
1
Intra-context conflicts in retrieval-augmented generation can be effectively resolved by a dual-confidence approach, leading to significant performance improvements in answering complex queries.
SeKV achieves a remarkable 5.9% performance boost in long-context LLM inference while slashing GPU memory usage by over half.
SproutRAG achieves a 6.1% boost in information efficiency by intelligently structuring document chunks without relying on external LLMs or lossy summarization.
Topic metadata can dramatically enhance retrieval efficiency in RAG systems, achieving over 8 times faster performance without sacrificing evidence quality.
Reducing visual token usage by 46% while improving performance shows that CUAs can leverage more historical data effectively without overwhelming compute budgets.