Search papers, labs, and topics across Lattice.
This paper introduces multi-byte prediction (MBP) for byte-level hierarchical language models, addressing the inference speed bottleneck associated with generating one byte at a time. By implementing a variable-length prediction window and a novel attention-masking scheme, MBP allows for parallel byte generation while maintaining causal integrity. The results demonstrate that MBP achieves a Pareto-optimal balance across various generative tasks, significantly enhancing throughput with minimal performance degradation.
Multi-byte prediction accelerates byte-level language model inference without sacrificing performance, achieving a breakthrough in generative task efficiency.
Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, we introduce multi-byte prediction (MBP), which generates multiple bytes in parallel, speeding up inference with minimal performance impact and no additional parameters. MBP builds on the popular multi-token prediction (MTP) paradigm with two crucial innovations. First, we introduce a variable-length prediction window that aligns with the latent tokens, or segments, of a hierarchical LM. Second, we implement a novel attention-masking scheme that enables parallel byte prediction without violating causality. We show that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.