Search papers, labs, and topics across Lattice.
6
0
12
6
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Relying on stale spatial memory can more than double an agent's failure rate, revealing a hidden safety risk in memory-augmented VLMs.
MemoryCPT achieves a superior cost-performance trade-off for LLM agents by intelligently managing memory without overwhelming downstream models with context.
Leveraging user history can cut clarification requests in coding assistants by identifying and resolving recurring ambiguities, leading to more efficient coding sessions.
Classic cache policies like LRU and LFU fail in semantic retrieval, while the new SOLAR framework achieves up to 75% improvement over FIFO by leveraging learning-augmented strategies.
ChartCynics outperforms state-of-the-art models by nearly 29% in accurately interpreting misleading charts, showcasing the power of specialized agentic workflows.