Search papers, labs, and topics across Lattice.
3
0
5
3
Prompt Coverage Adequacy uncovers over 30% more faults than traditional code coverage, revolutionizing how we test LLM-generated code.
Today's agents are surprisingly bad at real-world terminal tasks, with even frontier models failing nearly 40% of the time on everyday workflows.
Hallucinating LLMs in enterprise workflows can be tamed by a new Hybrid Utility Minimum Bayes Risk (HUMBR) framework that synthesizes semantic and lexical signals to achieve consensus without ground truth.