Search papers, labs, and topics across Lattice.
3
0
10
6
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Leveraging user history can cut clarification requests in coding assistants by identifying and resolving recurring ambiguities, leading to more efficient coding sessions.
ChartCynics outperforms state-of-the-art models by nearly 29% in accurately interpreting misleading charts, showcasing the power of specialized agentic workflows.