Search papers, labs, and topics across Lattice.
4
0
5
4
HiLS-Attention achieves over 64x context length extrapolation with 90% retrieval accuracy, outperforming traditional full attention mechanisms.
The choice of tree traversal method can significantly alter the performance of Transformer Grammars, revealing trade-offs that could redefine how we approach syntactic modeling in NLP.
Probabilistic Transformers can now scale to 0.4B parameters and beat standard Transformers of the same size, thanks to a hyperparameter transfer trick.
YOCO++ proves you can halve the KV cache size in LLMs and still beat a standard Transformer, thanks to a clever residual connection trick.