Search papers, labs, and topics across Lattice.
This paper introduces DunphyBench, a benchmark designed to evaluate agents on long-horizon, human-centered embodied decision-making, emphasizing the integration of multimodal inputs to align with complex human preferences. The authors identify significant performance gaps between current agents and human decision-making, attributing part of this disparity to ineffective memory management in state-of-the-art VLM-driven agents. To address this, they propose MeMento, a preference-conditioned multimodal memory compressor that enhances decision-making accuracy by 7.18% while drastically reducing memory usage by 85.38%.
Agents can now make more accurate decisions by effectively compressing multimodal memory, closing the gap with human performance in complex environments.
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents to accumulate evidence over long horizons, interpret implicit user preferences, and compare multiple candidates under partial observations. In this work, we propose DunphyBench, a new benchmark for evaluating agents on long-horizon human-centered embodied decision-making, where the agent must navigate through multiple embodied housing environments and make decisions that align with multi-dimensional human preferences. Unlike standard embodied reasoning tasks that often focus on procedural planning or immediate goal completion, our setting requires agents to integrate multimodal, multi-source input into coherent knowledge that supports complex reasoning across long horizon. The evaluation results reveal that there is a substantial gap between current agents and human performance. Furthermore, our diagnosis of state-of-the-art VLM-driven agents reveals that memory management is one of the bottlenecks, where raw multimodal history introduces noise that hinders decision quality. Motivated by this finding, we design MeMento, a preference-conditioned multimodal memory compressor that selectively compresses decision-relevant information from long-horizon history based on user preferences with a fixed set of memory tokens. Experiments show that MeMento helps VLM-driven agents improve accuracy by 7.18%, while reducing memory usage by 85.38% compared to the strongest baseline.