Search papers, labs, and topics across Lattice.
3
0
9
2
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Relying on stale spatial memory can more than double an agent's failure rate, revealing a hidden safety risk in memory-augmented VLMs.
ChartCynics outperforms state-of-the-art models by nearly 29% in accurately interpreting misleading charts, showcasing the power of specialized agentic workflows.