Search papers, labs, and topics across Lattice.
5
1
7
2
It is shown that data-agnostic methods, such as parameter averaging and dynamic selection, often fail to combine knowledge from logically disjoint fine-tuning datasets, and that LoRA reuse relies more on shallow pattern matching than on logical integration of existing knowledge.
Transformers can length-generalize on specific regular languages, revealing a hidden algebraic property that classical finite decomposition theory fails to capture.
Transforming our understanding of Transformers, this work reveals that learnability may be as crucial as expressivity in optimizing large language models.
LLM-based peer review systems can be made significantly more robust against adversarial manipulation via a co-evolutionary GAN approach that anticipates novel attacks.
Chain-of-Thought reasoning in Transformers hits a surprising expressivity ceiling when generalizing to longer sequences, unless you let your vocabulary grow with the problem size and use "signpost" tokens.