Search papers, labs, and topics across Lattice.
3
3
7
2
CrEST redefines credit assignment in RL by shifting the teacher's role from directing updates to modulating their magnitude, leading to substantial performance improvements in multi-turn agent training.
MLLMs still struggle with real-world document understanding, but a new benchmark and reinforcement learning approach can significantly improve their ability to extract structured information from receipts.
Forget synthetic data and overfitting: Environment Tuning lets LLM agents learn complex tool-use behaviors directly from the environment, slashing data needs and boosting generalization.