Search papers, labs, and topics across Lattice.
Thanks:
13
0
13
5
VoxAudio revolutionizes vocalized audio synthesis by embedding intelligible speech seamlessly within complex soundscapes, outperforming traditional methods that compromise on clarity and control.
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
HierCAD achieves unprecedented fidelity in CAD generation by aligning structural reasoning with geometric parameters, setting a new benchmark for text-to-CAD systems.
NaviCache redefines how we approach computational efficiency in video generation, achieving superior error judgment and performance without the burdens of traditional calibration methods.
ForeAgent outperforms existing deepfake detection methods by 16.41% while continuously evolving its reasoning capabilities through self-reflection and high-quality sample generation.
A unified taxonomy of audio editing tasks reveals the transformative potential of foundation models in reshaping how we interact with sound.
Full-duplex dialogue systems are often mischaracterized, with many claiming capabilities they cannot deliver due to training limitations.
Spatial-Omni achieves superior spatial audio understanding by seamlessly integrating FOA encoding into existing LLMs, outperforming traditional models without compromising general audio processing.
SwanSphere achieves real-time, high-fidelity spatial audio generation from panoramic video and text, overcoming the latency and spatial accuracy limitations of existing methods.
Forget OCR: DocRetriever achieves state-of-the-art multimodal document retrieval by cleverly combining layout-aware sparse embeddings with a reasoning-augmented reranker.
Current speech generation models still fall short in maintaining consistency and capturing nuanced expressiveness when generating long-form speech, despite advances in high-fidelity synthesis.
Persona-agnostic memory systems cause role-playing agents to forget who they are, but a dual-memory system can help them remember.
Current reward models for spoken dialogue systems are missing crucial paralinguistic and natural speech elements, but this new model closes the gap by operating directly on speech and outperforming existing audio LLMs.