Search papers, labs, and topics across Lattice.
8
0
10
11
This work proposes Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory, and studies two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances.
Achieving real-time ASR performance on edge devices, VibeVoice-ASR-BitNet outpaces Whisper.cpp by up to 2.3x while only modestly sacrificing accuracy.
Experiential Learning outperforms traditional reinforcement learning by providing richer feedback, leading to better generalization and reduced reward hacking in LLM training.
ReOPD turns the costly process of agent-environment interaction into a reusable offline resource, achieving faster and more efficient multi-turn distillation.
Language models can learn directly from real-world user interactions, boosting performance without human annotations or simulated environments.
By rethinking RLHF, MicroCoder-GRPO enables smaller code generation models to rival larger counterparts, achieving significant performance gains and revealing 34 training insights.
Unlock 33% faster LLM inference on commodity GPUs with SlideSparse, which finally brings hardware-accelerated (2N-2):2N sparsity to the masses, bridging the accuracy gap left by NVIDIA's strict 2:4 pruning.
1.58-bit LLMs are surprisingly more resilient to sparsity than their full-precision counterparts, opening new avenues for extreme compression.