Search papers, labs, and topics across Lattice.
6
0
9
7
Selecting relevant evidence spans for test-time training can boost long-context LLM accuracy by up to 15%.
Achieving near-autoregressive accuracy while boosting decoding speed by over 2.4 times could redefine efficiency benchmarks in generative reasoning tasks.
Re-ranking can make or break user engagement, and GR2 boosts performance by over 18% by harnessing the power of LLMs in ways previously unexplored.
Forget specialized architectures: StepAudio 2.5 proves a single audio-language foundation, shaped by RLHF, can dominate ASR, TTS, and real-time dialogue simultaneously.
RLVR, the dominant training paradigm for audio language models, may be turning them into unfeeling "answering machines" that excel on benchmarks but fail the vibe check.
Highlighting pivotal evidence can boost LLM performance without altering the original context, leading to substantial improvements in reasoning tasks.