Search papers, labs, and topics across Lattice.
This survey systematically reviews Instruction-based Image Editing (IIE), detailing its evolution from GANs to diffusion and autoregressive models while categorizing editing tasks and methodologies for training data construction. It highlights the significance of standardized evaluation metrics and introduces the Comprehensive, in-Depth, and Diagnostic benchmark for IIE tasks (CDD-IIE Bench) to rigorously assess model performance. Key findings reveal critical technological milestones and comparative capabilities of existing open-source solutions, paving the way for future advancements in the field.
The introduction of the CDD-IIE Bench could redefine how we evaluate and compare instruction-based image editing systems.
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical “one-sentence image editing” systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.