Search papers, labs, and topics across Lattice.
3
0
6
0
TACO effectively mitigates the reinforcement of erroneous reasoning in LLMs by distinguishing between useful and unreliable tokens, leading to improved training stability and performance.
Nexus Sampling retains crucial tokens during KV cache eviction, achieving near-dense attention performance with dramatically reduced memory usage.
Solve SMoE load balancing at inference time without retraining by replicating heavily used experts and quantizing underutilized ones, achieving up to 1.4x imbalance reduction.