Search papers, labs, and topics across Lattice.
Shanghai Artificial Intelligence Laboratory
4
0
6
1
Visual tool-use in multimodal LLMs may create an illusion of effectiveness, with many models achieving gains that are not causally justified.
Gradient Immunity can significantly hinder malicious fine-tuning efforts, keeping attack success rates at pre-release levels while enhancing safety without user intervention.
Evolving adversarial strategies can achieve nearly 100% success against state-of-the-art LLMs while maintaining robust defenses that adapt in real-time.
DataShield reveals that aligning consensus subspaces across multiple LLMs can drastically enhance safety by filtering out risky fine-tuning data more effectively than previous methods.