Search papers, labs, and topics across Lattice.
15
0
18
0
Long-tail semantic failures in document understanding are exposed when models are forced to reason with a unified vocabulary of visual anchors rather than treating elements in isolation.
Surgeons can now rely on a system that reduces cognitive workload by predicting relevant surgical regions in real-time, transforming laparoscopic visualization.
Harmful videos paired with benign queries can exploit a critical vulnerability in Video LLMs, leading to nearly half of all attacks succeeding despite high content recognition accuracy.
Audio-Zero reveals that fine-grained auditory reasoning can be achieved without any external labels, transforming how we approach audio model training.
Achieving stable and accurate alignment of generative 3D models with partial observations could redefine the standards for 3D reconstruction in computer vision.
DICE transforms long-document retrieval by effectively preserving critical information from chunks, achieving up to a 60% increase in retrieval accuracy for documents over 4,000 tokens.
SelectStream reveals that selective memory allocation can dramatically enhance streaming video understanding, outperforming traditional methods by preserving scene perception without sacrificing performance.
Rationale-based fine-tuning may actually undermine clinical prediction accuracy, challenging the belief that teaching models "why" can enhance their performance.
LLMs can now reliably generate complex, editable SVG diagrams by explicitly optimizing for geometric constraints via reinforcement learning, opening the door to automated technical illustration.
TwinGate stops jailbreaks by tracking malicious intent across anonymized, interleaved queries with minimal overhead, something previous defenses couldn't do.
Energy-dissipation principles can revolutionize how we infer potential functions in noisy, incomplete data environments, achieving remarkable robustness in generalized diffusion processes.
LLMs get *more* creative at generating molecules when you add *more* constraints, defying the intuition that creativity thrives on freedom.
Forget flattening: VideoStir's spatio-temporal graph retrieval and intent-aware scoring unlocks more effective reasoning over long videos.
Prompt highlighting in LLMs gets a serious upgrade: PRISM-$\Delta$ steers models to focus on relevant text spans with better accuracy and fluency, even in long contexts.
Forget expensive retraining: PromptCD unlocks significant improvements in LLM alignment and VLM visual grounding simply by contrasting model responses to cleverly designed prompts at test time.