Search papers, labs, and topics across Lattice.
14
23
14
22
Experiential Learning outperforms traditional reinforcement learning by providing richer feedback, leading to better generalization and reduced reward hacking in LLM training.
ReOPD transforms costly multi-turn interactions into a reusable offline resource, achieving up to 4× faster rollouts while preserving accuracy.
Achieving comparable performance to full-precision models, BITEMBED slashes storage costs and enhances embedding efficiency with extreme low-bit quantization.
G2PO redefines agent actions and leverages a global state-transition graph, leading to a 22.2% boost in success rates for long-horizon tasks.
LLMs struggle with Office automation, scoring only 36.6% on a standardized proficiency exam, revealing a critical gap in their capabilities.
Achieving up to 7.6x faster decoding and 17.1x greater throughput, CLSA redefines efficiency in long-context LLMs without compromising accuracy.
Language models can learn directly from real-world user interactions, boosting performance without human annotations or simulated environments.
Forget massive datasets – targeted training on a smaller, carefully curated dataset of challenging competitive programming problems yields 3x faster gains in code generation performance.
By rethinking RLHF, MicroCoder-GRPO enables smaller code generation models to rival larger counterparts, achieving significant performance gains and revealing 34 training insights.
Forget unimodal tasks—UniM throws down the gauntlet for truly unified multimodal AI, demanding models juggle any combination of text, image, audio, video, code, documents, and 3D inputs and outputs in a single, interleaved stream.
Unlock 33% faster LLM inference on commodity GPUs with SlideSparse, which finally brings hardware-accelerated (2N-2):2N sparsity to the masses, bridging the accuracy gap left by NVIDIA's strict 2:4 pruning.
1.58-bit LLMs are surprisingly more resilient to sparsity than their full-precision counterparts, opening new avenues for extreme compression.
Speculative decoding gets a throughput boost of up to 4.32x by using reinforcement learning to dynamically balance drafting and verification.
A 1-bit LLM can match the performance of full-precision models, promising huge gains in efficiency.