Search papers, labs, and topics across Lattice.
Fudan University Shanghai Innovation Institute
6
0
10
4
Predicting the next token's KV entries can boost long-context LLM throughput by over 2.5 times without sacrificing latency or quality.
Only a subset of design interactions in heterogeneous LLM inference are binding constraints, revealing critical insights for optimizing deployment strategies.
daVinci-kernel outperforms the best existing RL-trained model in GPU kernel optimization by effectively co-evolving skill selection and execution strategies.
Pretraining isn't just about scaling data volume; daVinci-LLM's ablations reveal that data processing depth, domain-specific strategies, and compositional balance are equally critical for unlocking LLM capabilities.
Forget toy datasets: OpenSWE delivers 45K+ real-world, executable Python environments for leveling up your SWE agent, and it's all open-sourced.
Subtracting the mean from activations unlocks stable FP4 training for LLMs, closing the performance gap with BF16 without complex spectral methods.