Search papers, labs, and topics across Lattice.
12
0
12
X2Streaming-TTS achieves true token-level synthesis with a median time to first audio token of just 15.8 ms, outperforming traditional pseudo-streaming models in quality and responsiveness.
Real-time turn-taking detection can be achieved with unprecedented accuracy and low latency using a novel dual-head modeling approach.
MMLMs can be made 99% safer against harmful multimodal inputs without sacrificing utility, thanks to a novel calibration approach.
Achieving high-fidelity 3D scene reconstruction from a single panorama could revolutionize the way we create and interact with virtual environments.
Task-conditioned foveated perception can drastically enhance the robustness and efficiency of robotic foundation models by aligning policy learning with relevant visual evidence.
LLM judges may misinterpret peer review quality, favoring superficial traits over genuine analytical depth, raising questions about their reliability in academic assessments.
DMuon slashes training time for large models, achieving up to 163x faster optimizer steps while maintaining the benefits of matrix-aware updates.
SPWM cuts computational costs and energy consumption in image restoration tasks while maintaining high image quality, showcasing the untapped potential of spiking neural networks.
MIXGUARD achieves robust privacy protection in split learning without sacrificing model utility, outperforming existing defenses against advanced data reconstruction attacks.
Uncovering critical security and privacy vulnerabilities in foundation-model-powered robots could redefine how we approach their deployment in real-world applications.
VOID shatters the semantic structure of images to thwart unauthorized mimicry without sacrificing visual quality, achieving a 223% boost in defense efficacy.
By redefining action learning around semantic events, WALL-WM achieves unprecedented generalization across tasks and environments, outperforming traditional models.