Score: 0

Learning an Image Editing Model without Image Editing Pairs

Published: October 16, 2025 | arXiv ID: 2510.14978v1

By: Nupur Kumari , Sheng-Yu Wang , Nanxuan Zhao and more

Potential Business Impact:

Teaches computers to edit pictures without examples.

Business Areas:

Photo Editing Content and Publishing, Media and Entertainment

Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Current workarounds use synthetic training pairs that leverage the zero-shot capabilities of existing models. However, this can propagate and magnify the artifacts of the pretrained model into the final trained model. In this work, we present a new training paradigm that eliminates the need for paired data entirely. Our approach directly optimizes a few-step diffusion model by unrolling it during training and leveraging feedback from vision-language models (VLMs). For each input and editing instruction, the VLM evaluates if an edit follows the instruction and preserves unchanged content, providing direct gradients for end-to-end optimization. To ensure visual fidelity, we incorporate distribution matching loss (DMD), which constrains generated images to remain within the image manifold learned by pretrained models. We evaluate our method on standard benchmarks and include an extensive ablation study. Without any paired data, our method performs on par with various image editing diffusion models trained on extensive supervised paired data, under the few-step setting. Given the same VLM as the reward model, we also outperform RL-based techniques like Flow-GRPO.

UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying

CV and Pattern Recognition

Changes pictures using words, no training needed.

5 Aug 2025 1

89%

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

CV and Pattern Recognition

Teaches computers to edit pictures better with clearer instructions.

5 May 2025 2

89%

PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

CV and Pattern Recognition

Teaches computers to edit pictures using examples.

9 Jun 2025 1

View PDF Login to Bookmark

Page Count

29 pages

Learning an Image Editing Model without Image Editing Pairs

Teaches computers to edit pictures without examples.

Technical Abstract

UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

PairEdit: Learning Semantic Variations for Exemplar-based Image Editing