Google ResearchCourant Institute of MathematicalHarvardNYUSchool of Engineering and AppliedMay 21, 2026arXiv:2605.22471

Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna

AI Summary

This paper analyzes the impact of graph tokenization choices on the expressivity of graph transformers, focusing on spectral, random-walk, and adjacency tokenizations. They prove that different tokenizations induce distinct depth regimes for realizing the same graph computation and demonstrate the impossibility of efficient conversion between tokenization families within limited-depth transformers. Theoretical results are validated with experiments showing task-specific preferences for different tokenizations and the benefits of combining complementary tokenizations.

Key Contribution

Graph transformers can be fundamentally limited by their tokenization strategy, as some tokenizations provably preclude efficient learning of structural representations realizable with other tokenizations.

Abstract

Transformers have become a central architecture for graph learning, but their application to graphs requires first choosing a tokenization: a graph-to-token map that determines which structural information is exposed at the input. In this work, we show that this choice is a fundamental component of transformer expressivity. We examine three tokenizations that serve as building blocks for many existing graph tokenizations: spectral, random-walk, and adjacency tokenizations. We prove that different tokenizations induce distinct depth regimes: the same graph computation may be realizable by a shallow transformer under one tokenization, while requiring substantially larger depth under another. For example, we prove that random-walk tokenization is lossy for any walk length, making it impossible in general to recover the graph from it, and that while spectral tokenization is lossless, it is ill-conditioned for local tasks. We further show that although both random-walk and spectral tokenizations are derived from adjacency information, it is impossible for a limited-depth transformer to convert between tokenization families in general. In particular, we establish lower bounds and impossibility results showing that unfavorable tokenizations may preclude the efficient recovery of more suitable structural representations. Finally, we complement our theory with controlled experiments on synthetic and real-world tasks, validating the predicted separations and showing that different tasks favor different structural views, and combining complementary tokenizations allows the transformer to leverage distinct signals from each representation.

Architecture Design (Transformers, SSMs, MoE)Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

Related Papers