論文の概要: Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
- arxiv url: http://arxiv.org/abs/2606.29814v1
- Date: Mon, 29 Jun 2026 05:48:41 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-30 18:07:16.063703
- Title: Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
- Title(参考訳): 中性子ラボ拡散画像:高分解能画像合成のためのマスク付き離散拡散の促進
- Authors: Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu, Aditya Grover, Jan Kautz, Pavlo Molchanov,
- Abstract要約: Nemotron-Labs-Diffusion-Imageは、テキスト・画像合成のための最先端のマスク付き離散拡散モデルである。
標準的なMDMには自己修正機能がない。
Nemotron-Labs-Diffusion-Imageにはトークン編集機構が組み込まれている。
- 参考スコア(独自算出の注目度): 93.7120829382642
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard MDMs lack self-correcting capability because discrete tokens cannot be modified once they are unmasked. Second, although increasing the vocabulary size of discrete image tokenizers improves reconstruction fidelity, it introduces optimization difficulties for generative modeling as the per-token training signal becomes increasingly sparse. To address the first challenge, Nemotron-Labs-Diffusion-Image incorporates a token-editing mechanism that enables the model to dynamically revise already-unmasked tokens during inference, similar to how a sculptor iteratively refines their work. To tackle the second challenge, we propose a Grouped Cross-Entropy (GCE) objective that assigns positive learning signals to tokens neighboring the ground truth in embedding space, thereby alleviating signal sparsity. To further improve training efficiency, we implement a custom fused operator for GCE that significantly reduces VRAM usage in large-vocabulary settings. Experimental results demonstrate that these innovations substantially improve both training efficiency and image fidelity of masked discrete image generators, achieving a score of 0.90 on GenEval, 86.9 on DPG and 10.76 of HPSv3.
- Abstract(参考訳): 我々は,高解像度テキスト・画像合成のための最先端のマスク付き離散拡散モデル(MDM)であるNemotron-Labs-Diffusion-Imageを提案する。
マスク画像生成に関する以前の研究と比較すると、Nemotron-Labs-Diffusion-Imageは2つの重要な課題に対処している。
第一に、画像全体にわたる遅延表現を徐々に洗練する連続拡散モデルとは異なり、標準的なMDMは、個別のトークンがマスキングされた後に修正できないため、自己修正能力が欠如している。
第二に、離散画像トークン化器の語彙サイズが大きくなると、再構成の忠実度が向上するが、個々の学習信号が疎結合になるにつれて、生成モデルの最適化が困難になる。
最初の課題に対処するため、Nemotron-Labs-Diffusion-Imageにはトークン編集機構が組み込まれている。
第2の課題に取り組むために、埋め込み空間における基底真理近傍のトークンに正の学習信号を割り当てるグループクロスエントロピー(GCE)の目標を提案する。
トレーニング効率をさらに向上するため,大語彙設定におけるVRAM使用量を大幅に削減する,GCE用カスタムフューズ演算子を実装した。
実験の結果、これらの革新によってマスク付き離散画像生成装置のトレーニング効率と画像忠実度が大幅に向上し、GenEvalでは0.90点、DPGでは86.9点、HPSv3では10.76点となった。
関連論文リスト
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation [52.261584726401686]
凍結した視覚基盤モデルの上に画像トークン化器を直接構築するための新しい方向を示す。
これらの設計に基づき,提案する画像トークン装置であるVFMTokは,画像再構成と生成品質の大幅な向上を実現している。
論文 参考訳(メタデータ) (2025-07-11T09:32:45Z) - Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models [92.18057318458528]
Token-ShuffleはTransformerにおける画像トークンの数を減らす新しい方法である。
我々の戦略は、事前訓練されたテキストエンコーダを必要とせず、MLLMが超高解像度画像合成をサポートできるようにする。
GenAIベンチマークでは、2.7Bモデルがハードプロンプトで0.77点、ARモデルLlamaGenが0.18点、拡散モデルLDMが0.15点である。
論文 参考訳(メタデータ) (2025-04-24T17:59:56Z) - Masked Autoencoders Are Effective Tokenizers for Diffusion Models [56.08109308294133]
MAETokは自己エンコーダであり、再構築の忠実さを維持しながら意味的にリッチな潜在空間を学習する。
MaETokは1.69のgFIDで76倍高速トレーニングが可能で、512x512世代で31倍高い推論スループットを実現している。
論文 参考訳(メタデータ) (2025-02-05T18:42:04Z) - Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis [62.57727062920458]
本稿では,非自己回帰型マスク画像モデリング(MIM)をSDXLのような最先端拡散モデルに匹敵するレベルまで高めるMeissonicを提案する。
高品質なトレーニングデータを活用し、人間の嗜好スコアから得られるマイクロ条件を統合し、特徴圧縮層を用いる。
我々のモデルは、高画質の高精細画像を生成する際に、SDXLのような既存のモデルに適合するだけでなく、しばしば性能を上回ります。
論文 参考訳(メタデータ) (2024-10-10T17:59:17Z) - Unified Auto-Encoding with Masked Diffusion [15.264296748357157]
我々はUMD(Unified Masked Diffusion)と呼ばれる,統合された自己監督的目標を提案する。
UMDは、パッチベースとノイズベースの破損テクニックを1つの自動エンコーディングフレームワークに組み合わせている。
下流の生成および表現学習タスクにおいて、高いパフォーマンスを達成する。
論文 参考訳(メタデータ) (2024-06-25T16:24:34Z) - Cross-view Masked Diffusion Transformers for Person Image Synthesis [21.242398582282522]
ポーズ誘導画像生成のための新しい拡散モデルであるX-MDPTを提案する。
X-MDPTは、潜伏パッチで動作するマスク付き拡散トランスフォーマーを用いて、自分自身を区別する。
我々のモデルはDeepFashionデータセットにおける最先端のアプローチよりも優れています。
論文 参考訳(メタデータ) (2024-02-02T15:57:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。