論文の概要: AdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing
- arxiv url: http://arxiv.org/abs/2603.21615v1
- Date: Mon, 23 Mar 2026 06:22:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-24 19:11:39.522902
- Title: AdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing
- Title(参考訳): AdaEdit:フローベース画像編集のための適応的時間・チャネル変調
- Authors: Guandong Li, Zhaobin Chu,
- Abstract要約: フローマッチングモデルにおけるインバージョンベースの画像編集は、トレーニング不要でテキスト誘導された画像操作のための強力なパラダイムとして登場した。
既存の方法は、注入要求の本質的に不均一な性質を無視した固定注入戦略でこの問題に対処する。
AdaEditは、このジレンマを2つの補完的な革新を通じて解決する、トレーニング不要な適応編集フレームワークである。
- 参考スコア(独自算出の注目度): 10.474377498273205
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Inversion-based image editing in flow matching models has emerged as a powerful paradigm for training-free, text-guided image manipulation. A central challenge in this paradigm is the injection dilemma: injecting source features during denoising preserves the background of the original image but simultaneously suppresses the model's ability to synthesize edited content. Existing methods address this with fixed injection strategies -- binary on/off temporal schedules, uniform spatial mixing ratios, and channel-agnostic latent perturbation -- that ignore the inherently heterogeneous nature of injection demand across both the temporal and channel dimensions. In this paper, we present AdaEdit, a training-free adaptive editing framework that resolves this dilemma through two complementary innovations. First, we propose a Progressive Injection Schedule that replaces hard binary cutoffs with continuous decay functions (sigmoid, cosine, or linear), enabling a smooth transition from source-feature preservation to target-feature generation and eliminating feature discontinuity artifacts. Second, we introduce Channel-Selective Latent Perturbation, which estimates per-channel importance based on the distributional gap between the inverted and random latents and applies differentiated perturbation strengths accordingly -- strongly perturbing edit-relevant channels while preserving structure-encoding channels. Extensive experiments on the PIE-Bench benchmark (700 images, 10 editing types) demonstrate that AdaEdit achieves an 8.7% reduction in LPIPS, a 2.6% improvement in SSIM, and a 2.3% improvement in PSNR over strong baselines, while maintaining competitive CLIP similarity. AdaEdit is fully plug-and-play and compatible with multiple ODE solvers including Euler, RF-Solver, and FireFlow. Code is available at https://github.com/leeguandong/AdaEdit
- Abstract(参考訳): フローマッチングモデルにおけるインバージョンベースの画像編集は、トレーニング不要でテキスト誘導された画像操作のための強力なパラダイムとして登場した。
このパラダイムの中心的な課題は、インジェクションジレンマ(インジェクションジレンマ)である。デノゲーション中にソース特徴を注入することは、元の画像の背景を保存するが、同時に編集されたコンテンツを合成するモデルの能力を抑圧する。
既存の手法では、時間的スケジュールのバイナリオン/オフ、均一な空間混合比、チャネルに依存しない潜在摂動といった、時間的およびチャネル的両方の次元にまたがるインジェクション要求の本質的に異質な性質を無視した、固定的なインジェクション戦略でこの問題に対処している。
本稿では、このジレンマを2つの補完的な革新を通じて解決する、トレーニング不要な適応編集フレームワークであるAdaEditを紹介する。
まず,ハードバイナリカットオフを連続的な減衰関数(シグモイド,コサイン,リニア)に置き換えるプログレッシブ・インジェクション・スケジュールを提案する。
第2に、チャネル選択遅延摂動(Channel-Selective Latent Perturbation)を導入し、これは、逆潜時とランダム潜時の間の分布ギャップに基づいてチャネルごとの重要度を推定し、構造エンコーディングチャネルを保ちながら、編集関連チャネルを強く摂動する。
PIE-Benchベンチマーク(700枚の画像、10種類の編集タイプ)の大規模な実験では、AdaEditはLPIPSの8.7%の削減、SSIMの2.6%の改善、PSNRの2.3%の改善を実現し、競合するCLIPの類似性を維持している。
AdaEditは完全にプラグアンドプレイで、Euler、RF-Solver、FireFlowなど複数のODEソルバと互換性がある。
コードはhttps://github.com/leeguandong/AdaEditで入手できる。
関連論文リスト
- Stable Flow: Vital Layers for Training-Free Image Editing [74.52248787189302]
拡散モデルはコンテンツ合成と編集の分野に革命をもたらした。
最近のモデルでは、従来のUNetアーキテクチャをDiffusion Transformer (DiT)に置き換えている。
画像形成に欠かせないDiT内の「硝子層」を自動同定する手法を提案する。
次に、実画像編集を可能にするために、フローモデルのための改良された画像反転手法を提案する。
論文 参考訳(メタデータ) (2024-11-21T18:59:51Z) - Taming Rectified Flow for Inversion and Editing [57.3742655030493]
FLUXやOpenSoraのような定流拡散変換器は、画像生成やビデオ生成の分野で優れた性能を発揮している。
その堅牢な生成能力にもかかわらず、これらのモデルは不正確さに悩まされることが多い。
本研究では,修正流の逆流過程における誤差を軽減し,インバージョン精度を効果的に向上する訓練自由サンプリング器RF-rを提案する。
論文 参考訳(メタデータ) (2024-11-07T14:29:02Z) - COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing [57.76170824395532]
ビデオ編集は新たな課題であり、現在のほとんどの手法では、ソースビデオを編集するために、事前訓練されたテキスト・トゥ・イメージ(T2I)拡散モデルを採用している。
我々は,高品質で一貫したビデオ編集を実現するために,COVE(Cor correspondingence-guided Video Editing)を提案する。
COVEは、追加のトレーニングや最適化を必要とせずに、事前訓練されたT2I拡散モデルにシームレスに統合することができる。
論文 参考訳(メタデータ) (2024-06-13T06:27:13Z) - Eliminating Contextual Prior Bias for Semantic Image Editing via
Dual-Cycle Diffusion [35.95513392917737]
Dual-Cycle Diffusionと呼ばれる新しいアプローチは、画像編集をガイドするアンバイアスマスクを生成する。
提案手法の有効性を実証し,D-CLIPスコアを0.272から0.283に改善した。
論文 参考訳(メタデータ) (2023-02-05T14:30:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。