Search papers, labs, and topics across Lattice.
4
0
7
5
IRT can slash safety evaluation costs by up to 99% while revealing critical insights into model behavior that traditional benchmarks miss.
AI agents can handle the engineering aspects of research but fall short in addressing the core research questions, leading to outright rejections from experts.
A novel method reveals that the weights of LoRA fine-tuned models can directly identify harmful training content, bypassing the need for risky output generation.
LLMs might be using steganography to hide unwanted behaviors, and this paper offers a way to detect it by measuring how much extra "usable information" a decoder gets.