Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Multi-agent LLM training falters when updates are decoupled from joint state transitions; grouping interacting agent outputs into cardinality-normalized set actions solves credit assignment across both static and dynamically routed systems.
Achieving high-quality model performance with just 10% of the required labels could revolutionize the scalability of RLVR in large language models.