Search papers, labs, and topics across Lattice.
This paper introduces Counterfactual Conditional Likelihood (CCL) rewards to address redundant exploration in multiagent systems by scoring each agent's unique contribution to team exploration. CCL rewards agents for observations that are informative with respect to the joint exploration of the team, rather than solely for individual novelty. Experiments in continuous multiagent domains demonstrate that CCL accelerates learning in sparse reward environments requiring tight coordination.
CCL rewards unlock faster learning in multiagent systems by rewarding agents for the unique information they contribute to the *team's* exploration, not just their own.
Efficient exploration is critical for multiagent systems to discover coordinated strategies, particularly in open-ended domains such as search and rescue or planetary surveying. However, when exploration is encouraged only at the individual agent level, it often leads to redundancy, as agents act without awareness of how their teammates are exploring. In this work, we introduce Counterfactual Conditional Likelihood (CCL) rewards, which score each agent's exploration by isolating its unique contribution to team exploration. Unlike prior methods that reward agents solely for the novelty of their individual observations, CCL emphasizes observations that are informative with respect to the joint exploration of the team. Experiments in continuous multiagent domains show that CCL rewards accelerate learning for domains with sparse team rewards, where most joint actions yield zero rewards, and are particularly effective in tasks that require tight coordination among agents.