Search papers, labs, and topics across Lattice.
5
0
7
A preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressive speech models that enables DPO-style preference learning without explicit sequence likelihoods while preserving direct supervision on preferred realizations is proposed.
Systems that excel in dialect identification often leverage unique acoustic features, while ASR performance hinges on data normalization strategies.