Search papers, labs, and topics across Lattice.
5
0
6
Poisoning pretraining data through public discussion interfaces poses a significant threat, with the potential for undetectable harmful behaviors in language models.
Automatic harness evolution may not be the silver bullet for LLM performance it was thought to be, often lagging behind simpler scaling methods.
Tmax sets a new standard for terminal agent performance with a surprisingly simple RL recipe that outshines larger models.
Language models can get a 12% boost in multi-turn conversation quality from just 10k examples of multi-turn training data, highlighting the critical gap between single-turn and multi-turn capabilities.
Forget passively analyzing model outputs – this new attack actively *trains* the model to regurgitate specific texts, revealing its training data with surprising accuracy.