Search papers, labs, and topics across Lattice.
10
0
12
4
Degrading image aesthetics can be a powerful strategy to thwart identity leakage in personalized visual content generated by diffusion models.
Early decision-making in multimodal reasoning can cut inference time without sacrificing accuracy, thanks to a novel dual evaluation of model competencies.
EvoEmbedding outperforms leading embedding models by adapting its representations in real-time, revolutionizing long-context retrieval and agentic memory integration.
Fine-tuning on the new OmniVideo-100K dataset boosts model performance by over 20% in audio-visual reasoning tasks, revealing the power of structured scripts in enhancing multimodal understanding.
Current autonomous AI agents are alarmingly unprepared for real-world adversarial attacks, often missing critical vulnerabilities in dynamic environments.
RAG systems can dynamically defend against retrieval corruption with a graph-based energy minimization approach, achieving superior robustness and response quality while slashing storage overhead.
Standard camera auto-exposure is blind to the needs of remote heart-rate monitoring, but a new method closes the gap to enable robust in-vehicle driver monitoring.
Current audio-language models are surprisingly bad at controlling and interpreting subtle vocal cues, failing in nearly half of situational dialogue scenarios.
Achieve high-fidelity, temporally coherent video editing without paired training data by combining sparse semantic control with dense motion and texture synthesis.
Current vision-language models stumble on subtle spine diagnoses, but a new dataset and benchmark expose these weaknesses and pave the way for clinically useful AI.