Search papers, labs, and topics across Lattice.
3
0
3
6
Decoupling compliance from learning enables multi-robot teams to excel in complex maritime missions while adhering to safety norms.
Why retrain from scratch? MAPEX extracts Pareto-optimal policies from existing single-objective RL agents, achieving comparable performance at a tiny fraction of the sample cost.
CCL rewards unlock faster learning in multiagent systems by rewarding agents for the unique information they contribute to the *team's* exploration, not just their own.