Search papers, labs, and topics across Lattice.
10
0
10
12
FreeToken transforms personal machines into powerful platforms for running massive AI models, enabling users to deploy frontier-scale intelligence without specialized infrastructure.
REAR transforms how we achieve user preference alignment in LLMs, enabling scalable realignment without costly retraining.
Achieving up to 88x efficiency gains, Taylor-Calibrate transforms the way we initialize hybrid linear attention models, drastically reducing the training burden.
Discrete diffusion policies, typically used for image generation, turn out to be surprisingly effective and efficient asynchronous executors for robots acting in dynamic environments, outperforming traditional continuous control methods.
Forget fancy quantization schemes – a simple token-wise INT4 quantization with Hadamard rotation is all you need to nearly match FP16 accuracy in LLM serving, without sacrificing throughput.
Sparse queries offer a surprisingly effective and efficient alternative to dense representations for image-to-3D generation, achieving comparable fidelity with less input-view bias.
Diffusion language models can now match autoregressive quality, thanks to a clever trick that forces them to agree with themselves.
Verifier-free evolution can now match or exceed the performance of verifier-based methods, while slashing API costs by 3x and boosting throughput by 10x, thanks to a clever model orchestration strategy.
K-means, previously relegated to offline processing, gets a 17.9x speed boost on modern GPUs thanks to Flash-KMeans' clever IO and contention optimizations.
Accelerate video generation by 45% without retraining, simply by pruning redundant latent patches and cleverly recovering attention scores.