Search papers, labs, and topics across Lattice.
5
0
8
AgentEval uncovers up to 38 hidden failure boundaries in conversational LLMs that traditional testing methods overlook.
Cleaning logs can drastically enhance the performance of model inference and anomaly detection by eliminating irrelevant noise, leading to more accurate insights from software systems.
MANGO reveals that automated oracle generation can match the diagnostic power of traditional methods, transforming how we test VLA-enabled robotic systems.
LLMs can now generate significantly better Java unit tests without mocks, achieving up to 25% higher branch coverage by learning from existing code and rigorously enforcing semantic constraints.
Hallucinations in RAG are far more pervasive than we thought: re-annotating existing benchmarks reveals 1.68x more instances of unsupported claims, and a new framework, RT4CHART, dramatically improves detection.