Search papers, labs, and topics across Lattice.
The University of Hong Kong
2
0
4
Neural Double Q-routing slashes mean completion times by up to 8.8% in complex industrial environments, outperforming traditional methods.
Train massive MoEs on Hopper GPUs faster and with less memory, even without native FP4 support, by cleverly quantizing activations and communication.