Search papers, labs, and topics across Lattice.
10
0
6
13
Amplifying just a handful of key neurons in the audio encoder can boost LALM accuracy on non-semantic speech attributes by over 25 points.
Compressing the KV cache of speech tokens can enhance decoding speed by over 1.49 times while improving performance on key benchmarks.
ORCA not only boosts performance by 26.4 points but also restores critical speaker identity cues that traditional models overlook.
Timestamp drift in ASR can be corrected with minimal parameter updates, achieving near-perfect alignment without sacrificing model performance.
CAAD achieves an 8% performance boost in speech language models while slashing inference latency and linguistic bias.
Voice recordings can reveal the oscillating states of Recurrent Respiratory Papillomatosis, providing a unique longitudinal perspective on a rare laryngeal disease.
Audio-Language models are cheating on benchmarks, acing tests even when they barely listen.
Text-only LLMs already contain surprisingly diverse levels of auditory knowledge, and this pre-existing knowledge strongly predicts their performance when adapted for audio-language tasks.
LALMs struggle to handle multiple concurrent audio inputs, but a simple input permutation strategy can significantly boost their multi-audio understanding without retraining.
Overcome LALM's struggles with localized dialectal prosody: a new Taiwanese audio-text dataset and fine-tuning strategy boosts accuracy by 6.5% on the TAU Benchmark.