論文の概要: Consistent-Inversion: Reverse Consistency Guidance for Structure-Preserving Visual Editing
- arxiv url: http://arxiv.org/abs/2606.07145v1
- Date: Fri, 05 Jun 2026 11:00:12 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-08 14:33:29.700967
- Title: Consistent-Inversion: Reverse Consistency Guidance for Structure-Preserving Visual Editing
- Title(参考訳): Consistent-Inversion:Reverse Consistency Guidance for Structure-Preserving Visual Editing (特集:情報ネットワーク)
- Abstract要約: Consistent-Inversionは、構造保存ビジュアル編集のためのトレーニング不要の逆整合ガイダンスフレームワークである。
SD3.5プロトコルを統一したプロトコルで、ターゲット・プロンプトアライメントを維持しながら、背景および構造的忠実性を改善する。
- 参考スコア(独自算出の注目度): 41.38183848746174
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Text-guided diffusion models have become effective tools for real-image visual editing, where the edited image must follow a target instruction while preserving editing-irrelevant structure. Most training-free editors rely on inversion: a source image is mapped to a noisy latent trajectory and the terminal latent is reused for target-prompt denoising. This reuse is useful for preservation, but it also couples source reconstruction and target editing. The resulting trajectory mismatch may either damage background/layout details or over-constrain the intended edit. This paper presents Consistent-Inversion, a training-free reverse consistency guidance framework for structure-preserving visual editing. Instead of treating the inverted source latent as a fixed initialization, Consistent-Inversion checks whether an intermediate target trajectory can be reversed toward the source inversion trajectory under the source prompt. To make this check well-defined, we construct an auxiliary target-side noise representation, perform source-guided reverse denoising, and use the resulting reverse consistency discrepancy as a correction signal for selected early target denoising steps. The method does not update model parameters, is compatible with inversion-based editors, and introduces only a small inference overhead when applied sparsely. Experiments on PIE-Bench show that Consistent-Inversion improves background and structural fidelity under a unified SD3.5 protocol while maintaining target-prompt alignment, and compatibility experiments further verify the same correction principle on classical Stable-Diffusion inversion pipelines.
- Abstract(参考訳): テキスト誘導拡散モデルはリアルイメージの視覚的編集に有効なツールとなり、編集不要な構造を維持しながら、編集された画像はターゲット命令に従う必要がある。
ほとんどのトレーニングフリーエディタは、インバージョンに依存しており、ソースイメージはノイズの多い潜在軌道にマッピングされ、端末ラテントはターゲットプロンプトの復調のために再利用される。
この再利用は保存に有用であるが、ソースの再構築とターゲット編集を兼ね備えている。
結果として得られた軌道ミスマッチは、バックグラウンド/レイアウトの詳細を傷つけるか、意図した編集を過剰に制限する。
本稿では,構造保存型視覚編集のためのトレーニング不要な逆整合ガイダンスフレームワークであるConsistent-Inversionを提案する。
Inverted source latent を固定初期化として扱う代わりに、Consistent-Inversion は、中間目標軌道がソースプロンプトの下でソース反転軌道へ逆転できるかどうかをチェックする。
このチェックを適切に定義するために、我々は、補助目標側ノイズ表現を構築し、ソース誘導逆復調を行い、結果の逆整合不一致を、選択した早期目標復調ステップの補正信号として利用する。
このメソッドはモデルパラメータを更新せず、インバージョンベースのエディタと互換性があり、スパースで適用された場合、わずかな推論オーバーヘッドしか導入しない。
PIE-Benchの実験では、コンシステント・インバージョン(Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Inversion, Consistent-Diffusion Inversion)により、SD3.5プロトコルの背景および構造フィリティが向上し、ターゲット・プロンプトアライメントを維持しながら向上することが示されている。
関連論文リスト
- ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition [60.25038184390257]
ReDesignは、編集可能な階層階層を成長させるエージェントフレームワークで、特定のツールを選択して、モダリティにまたがって構成する。
大規模な編集性を評価するために、909個のFigmaファイルと14,796個の編集命令からなるFigma Edit Replay Benchmarkを導入する。
論文 参考訳(メタデータ) (2026-07-28T10:53:33Z) - DuET: Dual Expert Trajectories for Diffusion Image Editing [42.49842620609682]
我々は、ソースイメージの条件付けを一時的に緩和する訓練不要推論手法であるDuETを紹介する。
DuETは命令の関連性、セマンティックな忠実さ、そして様々なモデルやベンチマークにおける知覚品質を継続的に改善する。
論文 参考訳(メタデータ) (2026-06-11T12:58:48Z) - Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints [7.503616785263929]
拡散に基づく点編集は、雑音潜時摂動の多様体に局所的な摂動を適用することで、画像の意味や細部を操作できる。
伝統的な点ベースの編集は、運動軌跡を定義するためにハンドルとターゲットポイントのペアに依存している。
中間編集ステップの評価とガイドを行うCLIPベースのモデルを導入し、生成した結果がセマンティックに一致し続けることを保証する。
論文 参考訳(メタデータ) (2026-05-13T11:05:31Z) - DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing [51.56484100374058]
我々は、事前訓練されたテキスト・ツー・イメージ(T2I)モデルのトレーニング不要な編集方法であるDirectEditを提案する。
DirectEditは、追加の神経機能評価(NFE)を導入することなく、固有の再構成エラーを除去する
実験により、DirectEditは効率よく正確な画像編集を実現し、最先端の手法よりも優れたパフォーマンスを提供することが示された。
論文 参考訳(メタデータ) (2026-05-04T10:09:18Z) - EditInfinity: Image Editing with Binary-Quantized Generative Models [64.05135380710749]
画像編集のためのバイナリ量子化生成モデルのパラメータ効率適応について検討する。
具体的には、画像編集のためのバイナリ量子化生成モデルであるEmphInfinityを適応させるEditInfinityを提案する。
テキストの修正と画像スタイルの保存を促進させる,効率的かつ効果的な画像反転機構を提案する。
論文 参考訳(メタデータ) (2025-10-23T05:06:24Z) - FlowCycle: Pursuing Cycle-Consistent Flows for Text-based Editing [12.424207508842192]
本研究では,新しいインバージョンフリーかつフローベース編集フレームワークであるFlowCycleを提案する。
本研究では,FlowCycleが最先端手法よりも優れた編集品質と一貫性を実現することを示す。
論文 参考訳(メタデータ) (2025-10-23T04:58:29Z) - FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing [47.908940130654535]
FlowAlignは、最適な制御ベースの軌道制御による一貫した画像編集のためのインバージョンフリーなフローベースフレームワークである。
我々の終点正規化は、編集プロンプトとのセマンティックアライメントのバランスと、軌道に沿ったソース画像との構造的整合性を示す。
FlowAlignは、ソース保存と編集の制御性の両方において、既存のメソッドよりも優れています。
論文 参考訳(メタデータ) (2025-05-29T06:33:16Z) - RIGID: Recurrent GAN Inversion and Editing of Real Face Videos [73.97520691413006]
GANのインバージョンは、実画像に強力な編集可能性を適用するのに不可欠である。
既存のビデオフレームを個別に反転させる手法は、時間の経過とともに望ましくない一貫性のない結果をもたらすことが多い。
我々は、textbfRecurrent vtextbfIdeo textbfGAN textbfInversion and etextbfDiting (RIGID) という統合されたリカレントフレームワークを提案する。
本フレームワークは,入力フレーム間の固有コヒーレンスをエンドツーエンドで学習する。
論文 参考訳(メタデータ) (2023-08-11T12:17:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。