Search papers, labs, and topics across Lattice.
3
0
6
IRT can slash safety evaluation costs by up to 99% while revealing critical insights into model behavior that traditional benchmarks miss.
A novel method reveals that the weights of LoRA fine-tuned models can directly identify harmful training content, bypassing the need for risky output generation.
RMCT reveals that you can reduce bias in language models without sacrificing their ability to articulate the very cues you're trying to mitigate.