Search papers, labs, and topics across Lattice.
4
0
7
EAPO revolutionizes LLM reasoning by dynamically integrating prior experiences, leading to consistent performance gains over traditional RLVR methods.
MOSS-Audio achieves state-of-the-art performance in audio understanding tasks by effectively integrating temporal cues and deep acoustic features, setting a new benchmark for audio-language models.
Rigid reward clipping throws away valuable information just beyond the boundary, but a simple stochastic rescue of these signals can substantially boost RLVR performance.
LLM benchmarks are riddled with hidden flaws that even human experts miss, but can be caught with an automated LLM auditor for under $15 per benchmark.