Search papers, labs, and topics across Lattice.
10
0
10
7
Achieving real-time ASR performance on edge devices, VibeVoice-ASR-BitNet outpaces Whisper.cpp by up to 2.3x while only modestly sacrificing accuracy.
Experiential Learning outperforms traditional reinforcement learning by providing richer feedback, leading to better generalization and reduced reward hacking in LLM training.
Despite high diagnostic accuracy, LLMs fail to choose valid recovery actions for over 60% of incidents, exposing a critical flaw in their operational utility.
Achieving comparable performance to full-precision models, BITEMBED slashes storage costs and enhances embedding efficiency with extreme low-bit quantization.
G2PO redefines agent actions and leverages a global state-transition graph, leading to a 22.2% boost in success rates for long-horizon tasks.
LLMs struggle with Office automation, scoring only 36.6% on a standardized proficiency exam, revealing a critical gap in their capabilities.
By cleverly combining YOCO's efficient attention with recursive computation, YOCO-U achieves a capability-efficiency sweet spot that neither technique can reach on its own.
Language models can learn directly from real-world user interactions, boosting performance without human annotations or simulated environments.
1.58-bit LLMs are surprisingly more resilient to sparsity than their full-precision counterparts, opening new avenues for extreme compression.
Unlock 33% faster LLM inference on commodity GPUs with SlideSparse, which finally brings hardware-accelerated (2N-2):2N sparsity to the masses, bridging the accuracy gap left by NVIDIA's strict 2:4 pruning.