Search papers, labs, and topics across Lattice.
Nanjing University
5
0
6
4
Achieving a 1.85x speedup in matrix multiplication on Ascend NPUs could redefine performance benchmarks for dynamic tensor operations.
Proxy models can slash RL post-training costs by up to 87.5% while maintaining critical fault reproduction capabilities.
MISA-T boosts rollout throughput by over 53% while preserving workload integrity, revolutionizing how RL pipelines manage heterogeneous demands.
Achieving nearly 5x speedup on RISC-V processors, RVANNS revolutionizes approximate nearest neighbor search by optimizing both vector representation and graph traversal.
Bole accelerates hybrid-attention LLMs by up to 4.72 times while slashing memory usage by up to 99 times, transforming the landscape of autoregressive decoding.