Search papers, labs, and topics across Lattice.
8
9
8
5
A novel simplification of RLHF is proposed from the perspective of variational inference, called V ariational A lignment with R e-weighting ( VAR), which transforms the alignment objective into an offline reward-driven re-weighted supervised fine-tuning (SFT) form.
Skill-SP not only pushes the performance ceiling of LLMs but also transforms initially misaligned models into high-performing agents through dynamic skill evolution.
Harness TTS achieves up to 35.6-point improvements in instruction-following win rates, revolutionizing expressive speech synthesis for voice assistants.
PolicyAlign enables LLMs to adapt to rapidly changing safety policies without relying on expensive supervision data, achieving significant safety improvements across diverse applications.
Conflict-aware filtering in GD^2PO boosts reinforcement learning efficiency by preventing negative signal cancellation from competing rewards.
Achieving high-dimensional audio representations without sacrificing generative quality, the F3-Tokenizer redefines the boundaries of audio autoencoding.
SeedPolicy overcomes the long-horizon limitations of Diffusion Policies in robot manipulation by compressing temporal information with a novel gated attention mechanism, achieving state-of-the-art imitation learning performance with significantly fewer parameters than vision-language-action models.
Ditch the RLHF complexity: a variational re-weighting approach turns alignment into stable, reward-driven SFT, rivaling existing methods.