Search papers, labs, and topics across Lattice.
This paper introduces EmoScope, a multi-agent framework that transforms emotional image editing by focusing on what can be edited rather than how to edit, utilizing emotion-conditioned affordance reasoning to identify image-specific editable spaces. By employing a semantic hierarchy to ensure content consistency and emotional expressiveness, EmoScope allows for interactive user refinement of editing strategies. In a large-scale evaluation, EmoScope significantly outperformed existing methods, with an 88.1% preference rate from participants across various emotional categories, highlighting its effectiveness in delivering context-sensitive edits.
EmoScope redefines emotional image editing by enabling users to discover unique, context-specific editing strategies rather than relying on static templates.
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.