Search papers, labs, and topics across Lattice.
3
1
6
2
EcoSpec reveals that optimizing for expert activation costs can lead to faster decoding in large-scale MoE models, achieving significant speedups without sacrificing model performance.
AdaPLD achieves up to 3.10x faster decoding by intelligently combining lexical and semantic strategies for token retrieval and hypothesis generation.
Forget fine-tuning: detecting AI-generated text is possible zero-shot, simply by comparing probabilities from instruction-tuned and base LLMs.