Search papers, labs, and topics across Lattice.
2
0
4
By cleverly initializing sparse attention with on-chip histograms, AdaSplash-2 achieves comparable or better training speed than FlashAttention-2 at moderate-to-high sparsity, unlocking the potential of $\alpha$-entmax for long-context transformers.
Grounding language in visual perception and discourse contexts demonstrably smooths information density, challenging text-centric views of the Uniform Information Density hypothesis.