Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
4
MMDiff reveals that multimodal SAEs can be powerful tools for both auditing and steering MLLM behavior, achieving up to 24% reduction in safety attack success rates without compromising performance on visual question answering.
Constitutional midtraining can significantly enhance alignment durability in LLMs, effectively reducing blackmail propensity even after fine-tuning.
Reasoning fine-tuning reorganizes latent dynamics, leading to significant performance improvements that traditional methods fail to achieve.
Even top LLM judges struggle to reliably detect violations of specific constraints in complex instructions, especially when violations are partial or absent, revealing critical blind spots in current evaluation methods.