Search papers, labs, and topics across Lattice.
3
0
6
RaBitQCache accelerates long-context LLM inference while cutting memory I/O by intelligently adapting token budgets based on attention sparsity.
Paraphrased prompts may seem equivalent, but they can drastically alter learning dynamics, with some prompts leading to better generalization and less forgetting.
GAIS enables the synthesis of complex tasks that not only outperform existing methods but also do so with dramatically improved data efficiency.