Search papers, labs, and topics across Lattice.
Email:
1
0
2
0
Recurrent linear attention does not compound quantization drift over long contexts—delta-rule updates and non-linear gates actively overwrite and compress noise, allowing an entire 27B hybrid model to run in NVFP4 W4A4 with near-zero quality loss up to 64K tokens.