Search papers, labs, and topics across Lattice.
This paper introduces a Coulomb particle model that optimizes the distribution of randomized features for kernel attention in Transformers, enhancing kernel-target alignment while employing a Riesz/Coulomb repulsive potential for regularization. The method demonstrates significant improvements in accuracy, calibration, and robustness across various benchmarks, all while maintaining the linear complexity of attention mechanisms. Notably, the approach allows for diverse and task-adaptive random features, which are crucial for effective learning in complex tasks.
Learning kernelized attention through a Coulomb particle model boosts performance without sacrificing the efficiency of linear attention.
Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean--Vlasov equation. We instantiate the method in linearized Transformer attention by learning positive random-feature maps in a first alignment phase, then freezing the kernel and training the remaining network parameters with cross-entropy. Experiments on synthetic classification and sentence-level benchmarks show that learned kernelized attention can improve accuracy, calibration, and robustness for several feature maps while preserving linear-attention inference complexity.