論文の概要: MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis
- arxiv url: http://arxiv.org/abs/2607.17634v1
- Date: Mon, 20 Jul 2026 07:40:56 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.538001
- Title: MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis
- Title(参考訳): MixDiffusion:MixDiffusion-based Uni-condition Text-to- Image Generation Models for Multi-condition Image Synthesis (特集:情報ネットワーク)
- Abstract要約: MixDiffusionはマルチ条件T2I生成のためのトレーニングフリー拡散フレームワークである。
理論的には、バウンディングボックス、キーポイント、スケッチ、深度マップ、参照画像、テキストなど、任意の数の制御条件をサポートしている。
- 参考スコア(独自算出の注目度): 46.56871839630665
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Recent advances in text-to-image (T2I) generation have enabled controllable image synthesis by incorporating conditions beyond text. However, most existing diffusion-based methods are limited to a single type of control condition (e.g., bounding boxes or keypoints), which restricts their flexibility. To address this limitation, we propose MixDiffusion, a training-free diffusion framework for multi-condition T2I generation. MixDiffusion theoretically supports an arbitrary number of control conditions, including bounding boxes, keypoints, sketches, depth maps, reference images, and text, by collaboratively integrating multiple pre-trained uni-condition diffusion models. The key insight of the proposed approach is to derive the predicted noise distribution in each denoising step of the diffusion-based multi-condition image generation model from the predicted noise distributions of multiple diffusion-based uni-condition models with a derived integration formula, which is supported by rigorous theory proof. Owing to its training-free nature, MixDiffusion is easy to deploy and readily extensible to new control modalities.
- Abstract(参考訳): テキスト・トゥ・イメージ(T2I)生成の最近の進歩により、テキスト以外の条件を取り入れることで、制御可能な画像合成が可能になった。
しかし、既存の拡散ベースのほとんどのメソッドは単一の種類の制御条件(例えば、バウンディングボックスやキーポイント)に制限されており、柔軟性が制限されている。
この制限に対処するため,マルチ条件T2I生成のためのトレーニングフリー拡散フレームワークであるMixDiffusionを提案する。
MixDiffusionは、有界ボックス、キーポイント、スケッチ、深度マップ、参照画像、テキストを含む任意の数の制御条件をサポートし、複数の事前訓練された一条件拡散モデルを統合する。
提案手法の主な洞察は,拡散型多条件画像生成モデルの各デノイングステップにおける予測ノイズ分布を,厳密な理論証明によって支持される導出積分式を持つ複数の拡散型一様条件モデルの予測ノイズ分布から導出することである。
トレーニングのない性質のため、MixDiffusionは簡単にデプロイでき、新しい制御モードに容易に拡張できる。
関連論文リスト
- Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion [60.186310080523135]
離散データ(テキスト)に対する自己回帰的アプローチと連続データ(画像)に対する拡散的アプローチへの生成的モデリングの分岐は、真に統一されたマルチモーダルシステムの開発を妨げる。
階層的二重プロセスとしてマルチモーダル生成を再構成する新しい確率的フレームワークである textbfCoM-DAD を提案する。
提案手法は、標準的なマスキングモデルよりも優れた安定性を示し、スケーラブルで統一されたテキスト画像生成のための新しいパラダイムを確立する。
論文 参考訳(メタデータ) (2026-01-07T16:21:19Z) - Constrained Discrete Diffusion [61.81569616239755]
本稿では,拡散過程における微分可能制約最適化の新たな統合であるCDD(Constrained Discrete Diffusion)を紹介する。
CDDは直接、離散拡散サンプリングプロセスに制約を課し、トレーニング不要で効果的なアプローチをもたらす。
論文 参考訳(メタデータ) (2025-03-12T19:48:12Z) - DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion [144.9653045465908]
拡散確率モデル(DDPM)に基づく新しい融合アルゴリズムを提案する。
近赤外可視画像融合と医用画像融合で有望な融合が得られた。
論文 参考訳(メタデータ) (2023-03-13T04:06:42Z) - ShiftDDPMs: Exploring Conditional Diffusion Models by Shifting Diffusion
Trajectories [144.03939123870416]
本稿では,前処理に条件を導入することで,新しい条件拡散モデルを提案する。
いくつかのシフト規則に基づいて各条件に対して排他的拡散軌跡を割り当てるために、余剰潜在空間を用いる。
我々は textbfShiftDDPMs と呼ぶメソッドを定式化し、既存のメソッドの統一的な視点を提供する。
論文 参考訳(メタデータ) (2023-02-05T12:48:21Z) - Unifying Diffusion Models' Latent Space, with Applications to
CycleDiffusion and Guidance [95.12230117950232]
関係領域で独立に訓練された2つの拡散モデルから共通潜時空間が現れることを示す。
テキスト・画像拡散モデルにCycleDiffusionを適用することで、大規模なテキスト・画像拡散モデルがゼロショット画像・画像拡散エディタとして使用できることを示す。
論文 参考訳(メタデータ) (2022-10-11T15:53:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。