Search papers, labs, and topics across Lattice.
This paper introduces a unified lookup-table inference method for ternary LLMs that addresses the inefficiencies in attention mechanisms caused by the mismatch between weight-dominated projections and K/V cache processing. By employing scaled multi-plane signed digits to represent runtime K/V states, the approach enables direct consumption of these digit planes by activation-derived tables, thus eliminating the need for dense K/V materialization. Experimental results demonstrate significant improvements in cache capacity, model quality, and hardware efficiency, validating the effectiveness of this method in optimizing ternary LLMs.
Ternary LLMs can now achieve efficient attention computation without the overhead of high-precision K/V processing, revolutionizing their performance.
Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision engine. Compressing this cache alone does not resolve the mismatch. To execute attention with the same lookup-table machinery as ternary projections, values accumulated in one reduction must retain a compatible representation and scale. This requirement also differs for keys and values during causal decoding, because newly generated values may belong to an unfinished cache block. This work develops a unified lookup-table inference approach for ternary LLMs. It stores runtime K/V states as scaled multi-plane signed digits organized around the reduction structure of attention. The resulting digit planes are consumed directly by activation-derived tables, avoiding dense K/V materialization between cache storage and attention computation. The design combines online K/V formation, bounded handling of incomplete value blocks, and a shared multi-stream datapath for Linear projections and attention. A constraint-guided search selects the representation and execution policy for a target quality--efficiency trade-off. Experiments on native and post-training ternary models validate the approach across cache capacity, model quality, and hardware efficiency.