Search papers, labs, and topics across Lattice.
8
0
13
0
On-demand safety interventions can significantly improve the performance of vision-language models without sacrificing their multimodal reasoning capabilities.
Video-OPSD reveals that focusing on privileged visual evidence can drastically improve the efficiency and effectiveness of self-distillation in Video-LLMs.
Language models can now be rigorously evaluated on their ability to generate falsifiable research ideas, not just stylistic fluency.
SafeCap boosts LVLM safety by up to 19 points through innovative caption-mediated reinforcement learning, outpacing traditional alignment methods.
Hard reflection can mask boundary-score errors in reflected diffusion, leading to misleading sample quality despite underlying inaccuracies.
Evolved playbooks can boost vulnerability detection rates by over 6x and outperform dedicated commercial products, reshaping the landscape of automated security auditing.
MLLMs still can't handle time-sensitive multimodal reasoning, often failing to integrate auditory and visual cues effectively in dynamic environments like a 4D escape room.
By incorporating language guidance into federated learning, SurgFed tackles the long-standing problem of tissue and task heterogeneity in surgical video understanding, leading to improved segmentation and depth estimation across diverse surgical settings.