Search papers, labs, and topics across Lattice.
Affiliation:
8
0
9
17
Parameter-space exploration can significantly enhance LLM reinforcement learning, yielding better performance with fewer training errors than traditional methods.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
ProReviewer outperforms larger models by up to 39% in peer review quality by enabling proactive investigation of research papers.
Structural differences in circuits may be misleading, as they often reflect interchangeable mechanisms rather than distinct functionalities.
Contextual prompts can significantly boost stance detection accuracy, but only if the right type of context is chosen鈥擫LM-generated descriptions shine while user metadata may backfire.
Current code reward models are myopic, mostly rewarding functional correctness, but Themis-RM learns to score code across multiple criteria and languages, opening the door to more nuanced and useful code generation.
Forget tedious multi-turn dialogues: Co-FactChecker's "trace-editing" lets human experts directly shape an LLM's reasoning process, leading to higher quality claim verification.
Most scientific claims in NLP die in obscurity, and even the survivors are more likely to be subtly reshaped than outright validated or debunked.