Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
MoNe slashes compute and memory costs by 80% for long-context inference while enabling Transformers to handle context lengths far beyond their original limits.
ReCo reweights GRPO to counteract the detrimental effects of response concentration, leading to significant improvements in reasoning performance on challenging benchmarks.