Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
7
Traditional probing methods fail to reveal the true memorization capabilities of large code LLMs, leading to inflated performance scores that obscure their genuine understanding.
Despite advances in LLMs, a staggering number of outputs still fail to meet structural and semantic requirements, revealing critical gaps in current generation methods.
Naive PDF parsing and chunking can severely bottleneck RAG performance on financial documents; careful selection yields substantial gains.
LLM-as-a-Judge accuracy hinges on temperature settings, revealing a task-dependent sweet spot that defies the common practice of fixed values like 0.1 or 1.0.