Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Multi-agent LLM training falters when updates are decoupled from joint state transitions; grouping interacting agent outputs into cardinality-normalized set actions solves credit assignment across both static and dynamically routed systems.
Decision-aware training signals outperform traditional next-observation predictions, leading to more effective learning in LLM agents.
Overcome simplicity bias in RL agents with PA-MoE, a mixture-of-experts architecture that learns task phases directly from the RL objective, leading to better expert specialization.