Search papers, labs, and topics across Lattice.
This paper introduces Relation, a novel token-mixing primitive that organizes pairwise evidence into Self and Exchange relations to enhance information flow in models. The authors demonstrate that their Full Relation approach consistently outperforms traditional Multi-Head Attention (MHA) in terms of validation negative log likelihood (NLL) across various model sizes, while FlashRelation significantly accelerates processing speed. Additionally, Hybrid Relation achieves competitive language modeling quality with a reduced layer configuration, suggesting a paradigm shift in token mixing strategies.
Full Relation not only outperforms MHA in validation NLL across all tested scales but also accelerates processing speed by up to 4.41 times with FlashRelation.
Attention directly derives normalized information flow from pairwise scores. We introduce Relation, an alternative token-mixing primitive that first organizes pairwise evidence into explicit Self and Exchange relations and derives information flow afterward. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, Hybrid Relation, and a KV-style Relation Cache. Across matched decoder-only models at approximately 10M, 30M, and 100M parameters, Full Relation achieves lower final validation NLL than MHA at all three scales. In a fixed-context reference benchmark, FlashRelation is 3.60-4.41x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. Hybrid Relation uses 75% Linear Relation layers and achieves strong language-modeling quality. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.