Feb 27, 2026arXiv:2602.23622

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

Shibo Hong, Boxian Ai, Jun Kuang, Wei Wang, FengJiao Chen, Zhongyuan Peng, Chenhao Huang, Yixin Cao

AI Summary

The paper introduces DeepLookEditBench (DLEBench), a new benchmark for evaluating instruction-based image editing models (IIEMs) specifically on their ability to edit small-scale objects (1-10% of image area). DLEBench comprises 1889 samples across seven instruction types, featuring challenges like occlusion and multi-object editing. They also propose a dual-mode evaluation framework and refined score rubrics to improve the robustness and alignment with human judgment when evaluating IIEMs on this task, revealing performance gaps in existing models.

Key Contribution

Instruction-based image editing models still struggle to edit small objects, with a new benchmark revealing significant performance gaps despite progress on existing benchmarks.

Abstract

Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate plausible adherence to instructions and strong reasoning ability on current benchmarks, their ability to edit small objects remains underexplored, despite its importance for precise local editing and refining details in both real and generated images. In this paper, we introduce DeepLookEditBench (DLEBench), the first benchmark dedicated to assessing the abilities of IIEMs in editing small-scale objects. Specifically, we construct a challenging testbed comprising 1889 samples across seven instruction types. In these samples, target objects occupy only 1%-10% of the image area, covering complex scenarios such as partial occlusion and multi-object editing. To ensure robust evaluation on this benchmark, we propose an evaluation protocol with refined score rubrics to minimize subjectivity and ambiguity in two criteria: Instruction Following and Visual Consistency. This protocol also introduces a dual-mode evaluation framework (Tool-driven and Oracle-guided Modes) addressing the misalignment between LMM-as-a-Judge and human judgements on DLEBench. Empirical results on 10 IIEMs reveal significant performance gaps in small-scale object editing, highlighting the need for specialized benchmarks to advance this ability.

Computer Vision Eval Frameworks & Benchmarks Multimodal Models

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

Related Papers