Search papers, labs, and topics across Lattice.
7
0
12
A preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressive speech models that enables DPO-style preference learning without explicit sequence likelihoods while preserving direct supervision on preferred realizations is proposed.
A lightweight Text Risk Score (TRS) is proposed, which estimates synthesis risk from interpretable text features without manual annotation or model training and shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.
It is shown that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling, are effective for segment-level SI evaluation.
Moderate quantum noise can actually enhance model performance by reducing complexity and generalization error, challenging conventional wisdom about noise in machine learning.
By intelligently pruning tokens based on spike timing and activation, Vision SmolMamba achieves state-of-the-art efficiency in spiking neural networks, outperforming even Spiking Mamba.