Search papers, labs, and topics across Lattice.
University of Science and Technology
12
0
10
3
HOMIE redefines video personalization by seamlessly integrating inter- and intra-subject inputs, achieving unprecedented fidelity and interaction accuracy.
Qwen-Music outperforms leading systems in musicality and audio quality, achieving state-of-the-art results across 13 of 16 metrics while generating songs from text and reinterpreting existing tracks.
UniSAE enables seamless, granular editing of speech attributes, allowing for precise control over speaker, emotion, and content in a unified framework.
Regularizing in the activation space with Sparse Autoencoders leads to superior continual learning performance in large language models, outperforming traditional weight-space methods.
Expressiveness preservation in speech-to-speech translation remains a significant hurdle, with systems scoring poorly on emotional and nonverbal fidelity despite achieving high translation accuracy.
AudioCALM achieves state-of-the-art performance in speech, sound, and music generation by seamlessly integrating diverse audio modalities into a single autoregressive framework.
STAR-VAE achieves state-of-the-art audio reconstruction fidelity by aligning latent space geometry with the hierarchical structure of audio signals.
SME-enabled optimizations can boost SEM performance by up to 6x on ARM CPUs, transforming how we approach wave propagation simulations.
Achieving high-fidelity audio generation with just four sampling steps, AudioX-Turbo dramatically cuts inference costs while enhancing performance across multimodal tasks.
Speech QA performance peaks at 4.17 Hz, challenging the assumption that higher frame rates always yield better reasoning outcomes.
Forget prompt engineering: MOSS lets autonomous agents rewrite their own source code to fix bugs and improve performance in production.
Disentangling high-level cross-modal reasoning from low-level modality-specific refinement in talking head generation yields superior lip-sync accuracy, video quality, and audio quality compared to entangled approaches.