論文の概要: An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing
- arxiv url: http://arxiv.org/abs/2606.15570v1
- Date: Sun, 14 Jun 2026 03:26:56 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-16 16:21:33.698247
- Title: An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing
- Title(参考訳): 単体・多体画像編集のための総合的ベンチマーク
- Authors: Yiwei Ma, Ke Ye, Weihuang Lin, Jiayi Ji, Xiaoshuai Sun, Tat-Seng Chua, Rongrong Ji,
- Abstract要約: 命令ベース画像編集(IIE)モデルの総合評価ベンチマークであるI2EBench2.0を提案する。
I2EBench2.0は、シングルラウンドとマルチラウンドの命令ベースの編集の両方を同時に評価する。
シングルラウンド評価には16次元、マルチラウンド評価には7次元が組み込まれている。
- 参考スコア(独自算出の注目度): 121.27321922764145
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In recent years, there have been notable advancements in the area of instruction-based image editing (IIE), which focuses on the automatic alteration of input images using a model. Nevertheless, assessing the effectiveness of these editing models poses a considerable challenge due to the intricate nature of instructions and the wide variety of edits. To tackle this problem, one urgent task in this domain is the development of a robust evaluation framework that can precisely gauge the quality of editing outcomes and offer valuable benchmarks to guide future improvements. To address this challenge, we present a comprehensive evaluation benchmark named I2EBench2.0, designed for single-round and multi-round assessment of IIE models. I2EBench2.0 has four key features: 1) Evaluation Across Single and Multi-rounds: I2EBench2.0 simultaneously evaluates both single-round and multi-round instruction-based edits, assessing the precision and consistency of the edits. 2) Extensive Evaluation Criteria: I2EBench2.0 encompasses a broad range of criteria, evaluating both high-level and low-level aspects of each IIE model. Specifically, it incorporates 16 dimensions for single-round evaluations and 7 for multi-round evaluations. 3) Alignment with Human Judgment: To ensure our benchmark aligns with human evaluation, we conducted a comprehensive user study for each criterion. 4) Research-driven Insights: By analyzing the strengths and weaknesses of current IIE models across all 16 single-round and 7 multi-round dimensions, we provide critical insights aimed at directing future research in this area. We tested eight recently developed IIE models using I2EBench2.0 and derived academic insights through meticulous comparison and analysis. The related code, dataset, and images generated by all IIE models are available on GitHub: https://github.com/cocoshe/I2EBench.
- Abstract(参考訳): 近年,モデルを用いた入力画像の自動修正に焦点を当てた命令ベース画像編集(IIE)が注目されている。
しかしながら、これらの編集モデルの有効性を評価することは、命令の複雑な性質と多種多様な編集が原因で大きな課題となる。
この問題に対処するために、この領域の緊急課題は、編集結果の質を正確に評価し、将来の改善を導くための貴重なベンチマークを提供する、堅牢な評価フレームワークを開発することである。
この課題に対処するため,I2EBench2.0 という総合評価ベンチマークを提示する。
I2EBench2.0には4つの重要な特徴がある。
I2EBench2.0はシングルラウンドとマルチラウンドの両方の命令ベースの編集を同時に評価し、編集の精度と一貫性を評価する。
2) 総合評価基準: I2EBench2.0 は,各 IIE モデルの高レベル・低レベル両面の評価において,幅広い基準を包含する。
具体的には、シングルラウンド評価に16次元、マルチラウンド評価に7次元を組み込む。
3) 人的判断とのアライメント: ベンチマークが人的評価と一致することを確認するため, 各基準について総合的なユーザスタディを実施した。
4)研究主導の視点:16の1ラウンドと7の多ラウンドのすべての領域において、現在のIIEモデルの長所と短所を分析することにより、この分野における将来の研究を導くことを目的とした重要な洞察を提供する。
I2EBench2.0を用いて,最近開発された8種類のIIEモデルを検証した。
関連するコード、データセット、およびすべてのIIEモデルによって生成されたイメージはGitHubで入手できる。
関連論文リスト
- Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models [12.603176617170504]
Omni IIE Benchは、実用的なアプリケーションシナリオにおいて、IIEモデルの編集一貫性を診断するために設計されたベンチマークである。
我々はOmni IIE Benchを用いた8つの主流IIEモデルの総合評価を行った。
本分析は,低セマンティックスケールから高セマンティックスケールタスクへの移行時のパフォーマンスギャップを初めて定量化する。
論文 参考訳(メタデータ) (2026-03-16T08:07:06Z) - DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model [10.609050605838805]
本稿では,IIEMの小型オブジェクト編集能力を評価するための最初のベンチマークであるDeepLookEditBenchを紹介する。
7つの命令タイプにわたる1889のサンプルからなる挑戦的なテストベッドを構築した。
これらのサンプルでは、ターゲットオブジェクトは画像領域の1%-10%しか占めておらず、部分閉塞や複数オブジェクト編集といった複雑なシナリオをカバーしている。
10個のIIEMの実証的な結果から、小規模オブジェクト編集における大きなパフォーマンスギャップが明らかとなり、この機能を前進させるための特別なベンチマークの必要性が浮かび上がっている。
論文 参考訳(メタデータ) (2026-02-27T02:59:34Z) - UniREditBench: A Unified Reasoning-based Image Editing Benchmark [52.54256348710893]
この研究は、推論に基づく画像編集評価のための統一ベンチマークUniREditBenchを提案する。
精巧にキュレートされた2,700個のサンプルからなり、8つの一次次元と18のサブ次元にわたる実世界シナリオとゲーム世界のシナリオをカバーしている。
このデータセットにBagelを微調整し、UniREdit-Bagelを開発した。
論文 参考訳(メタデータ) (2025-11-03T07:24:57Z) - Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing [84.16442052968615]
RISEBenchはReasoning-Informed ViSual Editing (RISE)の最初のベンチマークである。
RISEBenchは、時間、因果、空間、論理的推論の4つの主要な推論カテゴリに焦点を当てている。
オープンソースモデルとプロプライエタリモデルの両方を含む,9つの目立った視覚編集モデルを評価する実験を行った。
論文 参考訳(メタデータ) (2025-04-03T17:59:56Z) - I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing [67.05794909694649]
I2EBenchはIIEモデルによって生成された編集画像の品質を評価するための総合的なベンチマークである。
I2EBenchは2000以上の編集用イメージと4,000以上の対応するオリジナルおよび多様な命令で構成されている。
我々はI2EBenchをオープンソースとして公開し、すべての命令、入力画像、人間のアノテーション、すべての評価方法からの編集画像、新しいIIEモデルからの結果を評価するためのシンプルなスクリプトを公開します。
論文 参考訳(メタデータ) (2024-08-26T11:08:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。