Search papers, labs, and topics across Lattice.
7
0
6
4
Gradient Immunity can significantly hinder malicious fine-tuning efforts, keeping attack success rates at pre-release levels while enhancing safety without user intervention.
Multi-item users can face significantly more powerful poisoning attacks, but a new method achieves robust estimation without sacrificing performance.
AgentSnare can absorb nearly 47% of an attacker's tool calls while ensuring that no real targets are compromised, showcasing a new frontier in adaptive cybersecurity defenses.
Evolving adversarial strategies can achieve nearly 100% success against state-of-the-art LLMs while maintaining robust defenses that adapt in real-time.
Frontier LLMs can be induced to generate biologically hazardous sequences, with attack success rates reaching up to 100%.
DataShield reveals that aligning consensus subspaces across multiple LLMs can drastically enhance safety by filtering out risky fine-tuning data more effectively than previous methods.
AgentDoG 1.5 proves you can achieve GPT-5.4-level agent safety with open-source models trained on just 1k samples, slashing deployment overhead by two orders of magnitude.