Search papers, labs, and topics across Lattice.
TileMix introduces a novel tile-centric mixed-precision attention kernel designed to optimize long-context prefill in large language models (LLMs) by enabling spatial precision routing over hardware-aligned score tiles. This method partitions the attention matrix and utilizes compact bitmasks to dynamically adjust precision during inference, allowing for both FP16 and INT8 computations while maintaining dense token connectivity. The results demonstrate that TileMix significantly enhances prefill throughput and recovers quality lost with uniform INT8 precision, establishing a controllable accuracy-efficiency frontier across various model families.
TileMix recovers long-context quality lost under uniform INT8 while boosting prefill throughput, redefining the efficiency landscape for LLM inference.
Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online-softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware-aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped-query attention, variable-length batches, and INT8 key/value caches. Across LongEval, LV-Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long-context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy-efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang-Ulrica/TileMix.