論文の概要: Representation Distribution Matching for One-Step Visual Generation
- arxiv url: http://arxiv.org/abs/2607.02375v1
- Date: Thu, 02 Jul 2026 16:15:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.91616
- Title: Representation Distribution Matching for One-Step Visual Generation
- Title(参考訳): ワンステップ視覚生成のための表現分布マッチング
- Authors: Lan Feng, Wuyang Li, Eloi Zablocki, Matthieu Cord, Alexandre Alahi,
- Abstract要約: 表現分布マッチング(Representation Distribution Matching, RDM)の設計空間の解明
2つの設計軸、分布がどのように比較され、それらの表現が比較されるかを特定し、それに沿って制御された研究によって3つの結果が得られた。
同じレシピは4ステップのFLUX.2[klein]を1ステップのジェネレータに後付けし、GenEvalの4ステップバージョンである0.826から0.794、PickScoreの22.76から22.58を90 H200 GPU時間で上回った。
- 参考スコア(独自算出の注目度): 107.84208192304455
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We elucidate the design space of Representation Distribution Matching (RDM), our name for the paradigm that trains a one-step image generator by matching generated and reference feature distributions under frozen pretrained encoders. We identify two design axes, how the distributions are compared and the representations they are compared in, and controlled studies along them yield three findings. First, the classical MMD, which could not train convincing generators a decade ago, becomes a strong and scalable objective once estimated right. Second, the generated batch is then the operative variable, with an optimum above 2048, far beyond customary batch sizes. Third, any single representation can be gamed, driven below the real score while images stay visibly fake, so we match against a balanced battery of encoders and evaluate with SW_r14, a Sliced-Wasserstein distance over 14 encoders that is independent of the training loss and resists gaming. Combining the preferred choices yields improved RDM (iRDM): it sets the one-step state of the art on ImageNet at SW_r14 1.30, corroborated by PickScore, a human-preference proxy our objective never optimizes, which prefers it over the prior best one-step generator on 71.2% of matched samples. The same recipe post-trains the four-step FLUX.2 [klein] into a one-step generator, surpassing the four-step version on GenEval, 0.826 to 0.794, and on PickScore, 22.76 to 22.58, in 90 H200 GPU-hours. Project page: https://alan-lanfeng.github.io/rdm/.
- Abstract(参考訳): 凍結事前学習エンコーダの下で生成した特徴分布と参照特徴分布を一致させて一段階画像生成を訓練するパラダイムの名称であるRepresentation Distribution Matching (RDM) の設計空間を解明する。
2つの設計軸、分布がどのように比較され、それらの表現が比較されるかを特定し、それに沿って制御された研究によって3つの結果が得られた。
まず、従来のMDDは10年前、説得力のある発電機を訓練できなかった。
第二に、生成されたバッチはオペレーショナル変数であり、2048以上の最適化は、通常のバッチサイズよりもはるかに大きい。
第3に、画像が視覚的に偽装されている間、実際のスコア以下で、任意の1つの表現をゲーム化できるので、バランスのとれたエンコーダのバッテリと照合し、トレーニング損失に依存せず、ゲームに抵抗する14個のエンコーダのスライデッド・ワッサーシュタイン距離SW_r14と評価する。
SW_r14 1.30でImageNetのワンステップステート・オブ・ザ・アートをセットし、PickScoreによって裏付けられ、私たちの目的は最適化されず、71.2%のマッチしたサンプルに対して、以前の最高のワンステップジェネレータよりも好まれる。
同じレシピは4ステップのFLUX.2[klein]を1ステップのジェネレータに後付けし、GenEvalの4ステップバージョンである0.826から0.794、PickScoreの22.76から22.58を90 H200 GPU時間で上回った。
プロジェクトページ: https://alan-lanfeng.github.io/rdm/。
関連論文リスト
- David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training [8.352666876052616]
Diff-Instruct* (DI*) は1ステップのテキスト・ツー・イメージ生成モデルのためのデータ効率のよいポストトレーニング手法である。
提案手法は,人的フィードバックからオンライン強化学習としてアライメントを行う。
我々の2.6B emphDI*-SDXL-1stepモデルは、50ステップのFLUX-devモデルより優れている。
論文 参考訳(メタデータ) (2024-10-28T10:26:19Z) - Improved Distribution Matching Distillation for Fast Image Synthesis [54.72356560597428]
この制限を解除し、MDDトレーニングを改善する一連の技術であるMDD2を紹介する。
まず、回帰損失と高価なデータセット構築の必要性を排除します。
第2に, GAN損失を蒸留工程に統合し, 生成した試料と実画像との識別を行う。
論文 参考訳(メタデータ) (2024-05-23T17:59:49Z) - One-step Diffusion with Distribution Matching Distillation [54.723565605974294]
本稿では,拡散モデルを1ステップ画像生成器に変換する手法である分散マッチング蒸留(DMD)を紹介する。
約KLの発散を最小化することにより,拡散モデルと分布レベルで一致した一段階画像生成装置を強制する。
提案手法は,イメージネット64x64では2.62 FID,ゼロショットCOCO-30kでは11.49 FIDに到達した。
論文 参考訳(メタデータ) (2023-11-30T18:59:20Z) - Improved DDIM Sampling with Moment Matching Gaussian Mixtures [1.450405446885067]
本稿では,Gaussian Mixture Model (GMM) を逆遷移演算子 (カーネル) として,DDIM(Denoising Diffusion Implicit Models) フレームワーク内で提案する。
我々は,GMMのパラメータを制約することにより,DDPMフォワードの1次と2次の中心モーメントを一致させる。
以上の結果から, GMMカーネルを使用すれば, サンプリングステップ数が少ない場合に, 生成したサンプルの品質が大幅に向上することが示唆された。
論文 参考訳(メタデータ) (2023-11-08T00:24:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。