Search papers, labs, and topics across Lattice.
4
3
7
2
CrEST redefines credit assignment in RL by shifting the teacher's role from directing updates to modulating their magnitude, leading to substantial performance improvements in multi-turn agent training.
RODS synthesizes new training data on-the-fly, enabling agents to maintain high performance with 20x fewer trajectories than traditional methods.
Achieve 2.6x faster autoregressive world model inference without retraining by caching and selectively reusing block-level residuals across generation chunks.
Forget synthetic data and overfitting: Environment Tuning lets LLM agents learn complex tool-use behaviors directly from the environment, slashing data needs and boosting generalization.