Search papers, labs, and topics across Lattice.
3
0
7
8
AgentEval uncovers up to 38 hidden failure boundaries in conversational LLMs that traditional testing methods overlook.
LLMs struggle to connect identified root causes to their causal paths, achieving only 61.5% success in grounding diagnoses despite a 76% identification rate.
ADR transforms the landscape of code task generation, enabling LLMs to tackle genuinely novel and challenging coding problems that enhance their performance.