Search papers, labs, and topics across Lattice.
9
0
12
3
This work identifies and formalizes this failure mode, which is term destructive resource preemption: obtaining the resources required for a requested task by terminating, overwriting, evicting, or degrading an incumbent task.
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Semantic Affine Consistency transforms how visual tokenizers align latent semantics, resulting in a 26% reduction in gFID and setting a new benchmark for diffusion models.
Evolving safety harnesses using trajectory data can reduce agent safety risks by over 3x while enhancing overall utility.
Vid2WAM achieves superior task generalization and data efficiency by leveraging video diffusion priors, even with minimal expert demonstrations.
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
Language models are increasingly doing their real work in the "invisible" latent space, not the tokens we see.
Current LLM safety evaluations miss the mark: ATBench reveals how risks in realistic, multi-step agent interactions emerge over time, challenging even the strongest models.