Search papers, labs, and topics across Lattice.
2
0
5
AutoTrace reveals that even state-of-the-art LLMs falter in causal reasoning, struggling to differentiate between vulnerable and safe code despite a robust new dataset.
Calibration without comprehension reveals that fine-tuning LLMs for vulnerability detection fails to enhance their underlying security reasoning, with models achieving only 52.1% detection accuracy.