Search papers, labs, and topics across Lattice.
2
0
4
0
It is shown that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates, so a machine-readable schema for specifying each experimental setting is introduced to improve reproducibility and make the scope of cross-system comparisons explicit.
The lack of comprehensive benchmarks for AI blue teams leaves SOCs vulnerable, and this paper lays the groundwork for rectifying that gap.