Search papers, labs, and topics across Lattice.
5
23
7
9
Achieving real-time ASR performance on edge devices, VibeVoice-ASR-BitNet outpaces Whisper.cpp by up to 2.3x while only modestly sacrificing accuracy.
Achieving comparable performance to full-precision models, BITEMBED slashes storage costs and enhances embedding efficiency with extreme low-bit quantization.
1.58-bit LLMs are surprisingly more resilient to sparsity than their full-precision counterparts, opening new avenues for extreme compression.
Unlock 33% faster LLM inference on commodity GPUs with SlideSparse, which finally brings hardware-accelerated (2N-2):2N sparsity to the masses, bridging the accuracy gap left by NVIDIA's strict 2:4 pruning.
A 1-bit LLM can match the performance of full-precision models, promising huge gains in efficiency.