論文の概要: CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
- arxiv url: http://arxiv.org/abs/2603.20741v1
- Date: Sat, 21 Mar 2026 10:00:35 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-24 19:11:39.065238
- Title: CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
- Title(参考訳): CTCal: クロスステップ自己校正によるテキスト・画像拡散モデルの再考
- Abstract要約: 我々は、テキストプロンプトと生成された画像の正確なアライメントを実現するために、CTCal(Cross-Timestep Self-Calibration)を導入する。
CTCalはモデルに依存しないため、既存のテキスト・画像拡散モデルにシームレスに統合することができる。
- 参考スコア(独自算出の注目度): 39.59945414053394
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Recent advancements in text-to-image synthesis have been largely propelled by diffusion-based models, yet achieving precise alignment between text prompts and generated images remains a persistent challenge. We find that this difficulty arises primarily from the limitations of conventional diffusion loss, which provides only implicit supervision for modeling fine-grained text-image correspondence. In this paper, we introduce Cross-Timestep Self-Calibration (CTCal), founded on the supporting observation that establishing accurate text-image alignment within diffusion models becomes progressively more difficult as the timestep increases. CTCal leverages the reliable text-image alignment (i.e., cross-attention maps) formed at smaller timesteps with less noise to calibrate the representation learning at larger timesteps with more noise, thereby providing explicit supervision during training. We further propose a timestep-aware adaptive weighting to achieve a harmonious integration of CTCal and diffusion loss. CTCal is model-agnostic and can be seamlessly integrated into existing text-to-image diffusion models, encompassing both diffusion-based (e.g., SD 2.1) and flow-based approaches (e.g., SD 3). Extensive experiments on T2I-Compbench++ and GenEval benchmarks demonstrate the effectiveness and generalizability of the proposed CTCal. Our code is available at https://github.com/xiefan-guo/ctcal.
- Abstract(参考訳): 近年のテキスト・画像合成の進歩は拡散モデルによって大きく促進されているが、テキスト・プロンプトと生成された画像の正確なアライメントを実現することは、まだ持続的な課題である。
この困難は, 従来の拡散損失の限界に起因し, 微粒なテキスト画像対応をモデル化するための暗黙の監督のみを提供する。
本稿では,拡散モデル内での正確なテキスト画像アライメントの確立は,時間経過が増加するにつれて徐々に困難になる,という観測結果に基づいて,CTCal(Cross-Timestep Self-Calibration)を導入する。
CTCalは、より小さな時間ステップで形成された信頼性の高いテキストイメージアライメント(すなわち、クロスアテンションマップ)を活用して、よりノイズの多い大きな時間ステップでの表現学習を校正し、トレーニング中に明確な監督を提供する。
さらに,CTCalと拡散損失の調和的な統合を実現するための時間段階適応重み付けを提案する。
CTCalはモデルに依存しず、既存のテキストと画像の拡散モデルにシームレスに統合することができ、拡散ベース(例:SD 2.1)とフローベースアプローチ(例:SD 3)の両方を含む。
T2I-Compbench++とGenEvalベンチマークの大規模な実験は、提案したCTCalの有効性と一般化性を実証している。
私たちのコードはhttps://github.com/xiefan-guo/ctcal.comから入手可能です。
関連論文リスト
- ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models [33.09645476860831]
本稿では,画像RAGTurboを提案する。
テキストプロンプトが与えられた場合、関連するテキストイメージペアをデータベースから取得し、それらを生成プロセスの条件付けに使用する。
実験の結果,提案手法は既存の手法と比較して遅延を伴わずに高忠実度画像を生成することがわかった。
論文 参考訳(メタデータ) (2026-02-13T05:59:57Z) - Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers [79.94246924019984]
マルチモーダル拡散変換器 (MM-DiT) はテキスト駆動型視覚生成において顕著な進歩を遂げている。
マルチモーダルインタラクションを動的に再バランスするパラメータ効率向上手法である textbfTemperature-Adjusted Cross-modal Attention (TACA) を提案する。
本研究は,テキスト・画像拡散モデルにおける意味的忠実度向上における相互注意のバランスの重要性を強調した。
論文 参考訳(メタデータ) (2025-06-09T17:54:04Z) - Fast constrained sampling in pre-trained diffusion models [80.99262780028015]
任意の制約下で高速で高品質な生成を可能にするアルゴリズムを提案する。
我々の手法は、最先端のトレーニングフリー推論手法に匹敵するか、超越した結果をもたらす。
論文 参考訳(メタデータ) (2024-10-24T14:52:38Z) - Improving Consistency in Diffusion Models for Image Super-Resolution [28.945663118445037]
拡散法における2種類の矛盾を観測する。
セマンティックとトレーニング-推論の組み合わせを扱うために、ConsisSRを導入します。
本手法は,既存拡散モデルにおける最先端性能を示す。
論文 参考訳(メタデータ) (2024-10-17T17:41:52Z) - Faster Diffusion via Temporal Attention Decomposition [77.90640748930178]
テキスト条件拡散モデルにおける推論における注意機構の役割について検討する。
我々は、時間的注意づけ(TGATE)として知られるトレーニング不要の手法を開発した。
TGATEは、スケジュールされた時間ステップで注意出力をキャッシュして再利用することで、効率的に画像を生成する。
論文 参考訳(メタデータ) (2024-04-03T13:44:41Z) - Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention
Regulation in Diffusion Models [23.786473791344395]
拡散モデルにおけるクロスアテンション層は、生成プロセス中に特定のトークンに不均等に集中する傾向がある。
本研究では,アテンションマップと入力テキストプロンプトを一致させるために,アテンション・レギュレーション(アテンション・レギュレーション)という,オン・ザ・フライの最適化手法を導入する。
実験結果から,本手法が他のベースラインより一貫して優れていることが示された。
論文 参考訳(メタデータ) (2024-03-11T02:18:27Z) - Debiasing Text-to-Image Diffusion Models [84.46750441518697]
学習ベースのテキスト・トゥ・イメージ(TTI)モデルは、さまざまなドメインで視覚コンテンツを生成する方法に革命をもたらした。
近年の研究では、現在最先端のTTIシステムに非無視的な社会的バイアスが存在することが示されている。
論文 参考訳(メタデータ) (2024-02-22T14:33:23Z) - Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing [49.800746112114375]
本稿では,テキスト・画像拡散モデルのための学習後量子化手法(プログレッシブ・アンド・リラクシング)を提案する。
我々は,安定拡散XLの量子化を初めて達成し,その性能を維持した。
論文 参考訳(メタデータ) (2023-11-10T09:10:09Z) - CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for
Low-Dose CT Denoising and Generalization [41.64072751889151]
低線量CT(LDCT)画像は光子飢餓と電子ノイズによりノイズやアーティファクトに悩まされる。
本稿では,低用量CT (LDCT) 用新しいCOntextual eRror-modulated gEneralized Diffusion Model(CoreDiff)を提案する。
論文 参考訳(メタデータ) (2023-04-04T14:13:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。