Search papers, labs, and topics across Lattice.
2
0
3
0
Traditional LLM benchmarks can mislead, as they often fail to distinguish between different reasoning capabilities, collapsing multiple policies into a single equivalence class.
LLMs struggle to effectively integrate and order evidence in attack chain reconstruction, with top models only succeeding 39.6% of the time on critical tasks.