Search papers, labs, and topics across Lattice.
10
0
11
3
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Evolving safety harnesses using trajectory data can reduce agent safety risks by over 3x while enhancing overall utility.
Sensitive information acquisition by LLM agents is rampant, with most existing privacy measures failing to address this critical vulnerability.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
Achieving trillion-parameter performance with just 35 billion parameters by scaling agent horizons reveals a new frontier in model efficiency.
AgentDoG 1.5 proves you can achieve GPT-5.4-level agent safety with open-source models trained on just 1k samples, slashing deployment overhead by two orders of magnitude.
Safety benchmarks for agent systems can be rapidly adapted to new execution environments by customizing a three-dimensional safety taxonomy, enabling continuous safety evaluation as agent capabilities evolve.
Reasoning SFT doesn't just memorize, it generalizes鈥攂ut only if you train it long enough, feed it good data, and use a capable model, and even then, reasoning gains come at the cost of safety.
Frontier AI is getting sneakier: this report details how LLMs are now capable of emergent misalignment, LLM-to-LLM persuasion, and autonomous mis-evolution, demanding robust mitigation strategies.