Search papers, labs, and topics across Lattice.
This paper introduces FUSE, an adaptive framework for Active Functional Affordance Grounding that enables embodied agents to explore scenes and identify objects based on their functional properties rather than mere identity. By integrating uncertainty-driven exploration with a learned planner, FUSE efficiently selects informative viewpoints, addressing the limitations of existing methods that rely on fixed perspectives. Experimental results demonstrate that FUSE not only outperforms previous approaches in grounding accuracy but also reduces computational demands by 1.33 times compared to traditional explicit exploration methods.
FUSE achieves superior functional grounding performance while cutting computation costs, revolutionizing how agents interact with their environments.
Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire observations that reveal discriminative functional evidence. Existing affordance grounding methods operate from fixed viewpoints and lack mechanisms for deciding where to look when functional cues are occluded or incomplete. We introduce Active Functional Affordance Grounding, a new task in which an agent sequentially explores a scene to identify and spatially ground an object satisfying a functional query. To address this problem, we propose FUSE, an adaptive semantic-geometric evidence acquisition framework that combines explicit uncertainty-driven exploration with a learned amortized planner to efficiently select informative viewpoints. We further introduce a Habitat-based benchmark for evaluating active functional grounding. Experiments show that FUSE achieves the highest observed non-oracle grounding performance while reducing computation by 1.33x relative to fully explicit exploration, and remains effective across multiple affordance knowledge sources.