Search papers, labs, and topics across Lattice.
2
0
4
7
Chimera achieves 7.3x compute efficiency over traditional models while enabling zero-shot extrapolation from short video clips to significantly longer sequences.
Multi-Head Attention Residuals achieve superior validation loss by allowing Transformers to leverage multiple attention heads, revealing that subspace disagreement is a key factor in model performance.