Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
1
MoNe slashes compute and memory costs by 80% for long-context inference while enabling Transformers to handle context lengths far beyond their original limits.
Shrinking LLM reasoning for mobile devices is now possible: LoRA adapters, RL-based budget forcing, and KV-cache tricks let Qwen2.5-7B reason efficiently on-device.