論文の概要: CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
- arxiv url: http://arxiv.org/abs/2608.14546v1
- Date: Fri, 14 Aug 2026 17:59:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-17 20:14:29.464685
- Title: CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
- Title(参考訳): CPI-Bench: リアルタイム画像編集のための総合的で実践的でインテリジェントなベンチマーク
- Abstract要約: CPI-Benchは現実世界の画像編集のベンチマークである。
CPI-BenchはCPI-General-Bench、CPI-Practical-Bench、CPI-Intelligent-Benchの3つのコアサブセットから構成される。
- 参考スコア(独自算出の注目度): 22.386044581413312
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.
- Abstract(参考訳): 画像編集モデルの急速な進歩と、様々な領域にまたがる広範な応用により、これらのモデル機能を現実のシナリオに直接展開する必要性が高まっている。
しかし、既存のベンチマークは単純な単一イメージのタスクに限られており、カバー範囲が限られており、様々なモデルのパフォーマンスを効果的に区別できない。
その結果、複雑なマルチイメージ編集、要求の高い推論命令、実用的なデプロイメント設定において、モデルパフォーマンスを確実に評価することができない。
これらの制約に対処するため、実世界の画像編集のための包括的・実践的・知的なベンチマークであるCPI-Benchを提案する。
CPI-General-Benchは、様々な編集タスクを包括的にカバーし、マルチイメージ編集評価の先駆者であるCPI-Practical-Bench、高周波リアルタイムアプリケーションシナリオにフォーカスしたCPI-Practical-Bench、高度に要求される推論ベースの編集機能の評価に特化したCPI-Intelligent-Benchである。
CPI-Benchに基づく主流画像編集モデルの評価結果は、CPI-Benchがモデル間の性能差を高めることを示す。
一般的な編集機能のギャップの包括的かつ信頼性の高い定量化、実際のデプロイメントの有効性、高度な推論ベースの編集を提供し、将来の画像編集モデルの最適化のための貴重なガイダンスを提供する。
重要な点として、我々のランキング分析では、CPI-BenchがArena Image Edit Leaderboardと最高の整列を達成しており、人間の評価者の好みや知覚的判断を忠実に捉え、現実世界のユーザー体験の堅牢なプロキシとして機能していることを示している。
関連論文リスト
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing [53.21868575804758]
命令ベースのビデオ編集(IVE)のための包括的かつ構造化されたベンチマークを導入する。
本ベンチマークでは,編集タスクを空間,時間,音声,参照ベースの編集など,複数のビデオ特化次元に分解する。
本研究では,4つの相補的な次元(精度,保存性,リアリズム,一貫性)から編集品質を評価する評価フレームワークを提案する。
論文 参考訳(メタデータ) (2026-08-05T16:58:48Z) - Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling [18.3502725927898]
画像編集と報酬モデリングのための統合評価スイートであるEdit- and Edit-Rewardを紹介する。
Edit-Rewardには6つの段階的な課題カテゴリにまたがる2,388のアノテーション付きインスタンスが含まれている。
構造的推論と慎重に設計されたルーリックに基づく細粒度多次元評価フレームワークを採用する。
論文 参考訳(メタデータ) (2026-05-13T06:33:54Z) - DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning [26.88130913151649]
画像差分キャプション(IDC)は、2つの画像の違いを正確に識別する言語記述を生成する。
DiffCap-Benchは10の異なるカテゴリをカバーする総合的なIDCベンチマークである。
また,人間の有意差分リストに基づくLCM-as-a-Judge評価プロトコルを提案する。
論文 参考訳(メタデータ) (2026-05-06T05:12:41Z) - GEditBench v2: A Human-Aligned Benchmark for General Image Editing [58.86807672117726]
GEditBench v2は、23のタスクにまたがる1200の現実世界のユーザクエリを備えた包括的なベンチマークである。
また、視覚的整合性を評価するためのオープンソースのペアワイドアセスメントモデルであるPVC-Judgeを提案する。
PVC-Judgeは、オープンソースモデルの最先端評価性能を達成し、平均してGPT-5.1を超えている。
論文 参考訳(メタデータ) (2026-03-30T15:08:32Z) - I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models [78.62380562116135]
既存の画像編集ベンチマークは、タスクの範囲が限られており、評価範囲が不十分であり、手動のアノテーションに大きく依存している。
画像間編集モデルの総合的なベンチマークである textbfI2I-Bench を提案する。
I2I-Benchを用いて、多数の主流画像編集モデルをベンチマークし、様々な次元にわたる編集モデル間のギャップとトレードオフを調査した。
論文 参考訳(メタデータ) (2025-12-04T10:44:07Z) - UniREditBench: A Unified Reasoning-based Image Editing Benchmark [52.54256348710893]
この研究は、推論に基づく画像編集評価のための統一ベンチマークUniREditBenchを提案する。
精巧にキュレートされた2,700個のサンプルからなり、8つの一次次元と18のサブ次元にわたる実世界シナリオとゲーム世界のシナリオをカバーしている。
このデータセットにBagelを微調整し、UniREdit-Bagelを開発した。
論文 参考訳(メタデータ) (2025-11-03T07:24:57Z) - KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models [88.58758610679762]
KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark) は、認知的なレンズを通してモデルを評価するための診断ベンチマークである。
本研究は,3つの基礎知識タイプ(実例,概念,手続き)にまたがる編集タスクを分類する。
詳細な評価を支援するため,人間の研究により知識ヒントによって強化され,校正された新しい知識プラウザビリティ指標を組み込んだプロトコルを提案する。
論文 参考訳(メタデータ) (2025-05-22T14:08:59Z) - CompBench: Benchmarking Complex Instruction-guided Image Editing [63.347846732450364]
CompBenchは複雑な命令誘導画像編集のための大規模なベンチマークである。
本稿では,タスクパイプラインを調整したMLLM-ヒューマン協調フレームワークを提案する。
編集意図を4つの重要な次元に分割する命令分離戦略を提案する。
論文 参考訳(メタデータ) (2025-05-18T02:30:52Z) - PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM [17.89238060470998]
拡散に基づく画像編集モデルを評価することは、生成AIの分野において重要な課題である。
我々のベンチマークであるPixLensは、編集品質と遅延表現の絡み合いを総合的に評価する。
論文 参考訳(メタデータ) (2024-10-08T06:05:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。