Search papers, labs, and topics across Lattice.
2
0
4
2
MMDiff reveals that multimodal SAEs can be powerful tools for both auditing and steering MLLM behavior, achieving up to 24% reduction in safety attack success rates without compromising performance on visual question answering.
Even top LLM judges struggle to reliably detect violations of specific constraints in complex instructions, especially when violations are partial or absent, revealing critical blind spots in current evaluation methods.