Search papers, labs, and topics across Lattice.
3
0
4
HyperDFlash achieves remarkable decoding speedups and accuracy improvements by aligning residual streams with the MHC architecture, outperforming existing methods.
Achieve significant speedups in diffusion model inference, without training, by adaptively selecting the best predictor for each token at each step based on a low-cost probe of the first layer.
Autoregressive video models can now generate 4-minute videos without retraining, thanks to a clever inference-time hack that fixes positional embedding bias and injects dynamic priors.