論文の概要: Depth as Time in One-Step Generative Models
- arxiv url: http://arxiv.org/abs/2610.03626v1
- Date: Fri, 02 Oct 2026 17:20:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.513164
- Title: Depth as Time in One-Step Generative Models
- Title(参考訳): 一段階生成モデルにおける時間としての深さ
- Abstract要約: モデル独自の出力ヘッドを用いて中間層を復号化することにより,多段階拡散による復号化計算を再現可能であることを示す。
本稿では, 層間計算をフローとして明示的に扱うと, 時間条件付きブロックを1つずつトレーニングして, 層間をまたがってdenoise-then-renoise動作を記述した。
- 参考スコア(独自算出の注目度): 40.73691747944504
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.
- Abstract(参考訳): 最近の1段階生成モデルの波は、蒸留または学習フローマップを通じて拡散の多段階軌跡を圧縮し、高品質な画像を生成するための屈折点に達している。
生成が1つの前方通過に圧縮されたとき、多段拡散の認知軌道はどうなるのか?
複数ステップの拡散がサンプリングステップにまたがるデノゲーション計算は、1つのフォワードパスの深さにわたって展開し、モデル自身の出力ヘッドで中間層をデコードすることで復元できる。
最も興味深いのは、この深度計算が、フローマップが解くように訓練された輸送タスクに依存することである。
最も驚くべきケースはMeanFlowで、単一のネットワーク評価内で、短いトランスポート間隔を探索することで、デノナイズとレノナイズの両方が明らかになる。
対照的に、ドリフトモデルのようなタイムインデクシングされた輸送タスクなしで訓練されたジェネレータは同じ深さのデノイングを示さない。
その結果, 深度劣化現象を示すモデルは, 階層計算によりより圧縮可能であることがわかった。 MeanFlow \texttt{SiT-L/2} モデルは, パラメータの16.6\times$を1つの時間条件ブロックに圧縮することができる。
本稿では, 階層計算をフローとして明示的に扱うと, 一つの時間条件ブロックをトレーニングし, パラメータの16.6\times$でMeanFlow \textt{SiT-L/2}モデルを圧縮することができることを示す。
これらの結果から,拡散の時間計算は1ステップ生成によって除去されるのではなく,ネットワーク深度で再編成されることが示唆された。
関連論文リスト
- ODE$_t$(ODE$_l$): Shortcutting the Time and Length in Diffusion and Flow Models for Faster Sampling [33.87434194582367]
本研究では,品質・複雑さのトレードオフを動的に制御できる相補的な方向について検討する。
我々は,フローマッチングトレーニング中に時間と長さの整合性項を用い,任意の時間ステップでサンプリングを行うことができる。
従来の技術と比較すると、CelebA-HQとImageNetのイメージ生成実験は、最も効率的なサンプリングモードで最大3$times$のレイテンシの低下を示している。
論文 参考訳(メタデータ) (2025-06-26T18:59:59Z) - FlowDPS: Flow-Driven Posterior Sampling for Inverse Problems [51.99765487172328]
逆問題解決のための後部サンプリングは,フローを用いて効果的に行うことができる。
Flow-Driven Posterior Smpling (FlowDPS) は最先端の代替手段よりも優れています。
論文 参考訳(メタデータ) (2025-03-11T07:56:14Z) - Fast constrained sampling in pre-trained diffusion models [80.99262780028015]
任意の制約下で高速で高品質な生成を可能にするアルゴリズムを提案する。
我々の手法は、最先端のトレーニングフリー推論手法に匹敵するか、超越した結果をもたらす。
論文 参考訳(メタデータ) (2024-10-24T14:52:38Z) - Cache Me if You Can: Accelerating Diffusion Models through Block Caching [67.54820800003375]
画像間の大規模なネットワークは、ランダムノイズから画像を反復的に洗練するために、何度も適用されなければならない。
ネットワーク内のレイヤの振る舞いを調査し,1) レイヤの出力が経時的にスムーズに変化すること,2) レイヤが異なる変更パターンを示すこと,3) ステップからステップへの変更が非常に小さいこと,などが分かる。
本稿では,各ブロックの時間経過変化に基づいて,キャッシュスケジュールを自動的に決定する手法を提案する。
論文 参考訳(メタデータ) (2023-12-06T00:51:38Z) - UDPM: Upsampling Diffusion Probabilistic Models [33.51145642279836]
拡散確率モデル(DDPM、Denoising Diffusion Probabilistic Models)は近年注目されている。
DDPMは逆プロセスを定義することによって複雑なデータ分布から高品質なサンプルを生成する。
生成逆数ネットワーク(GAN)とは異なり、拡散モデルの潜伏空間は解釈できない。
本研究では,デノナイズ拡散過程をUDPM(Upsampling Diffusion Probabilistic Model)に一般化することを提案する。
論文 参考訳(メタデータ) (2023-05-25T17:25:14Z) - Fast Sampling of Diffusion Models via Operator Learning [74.37531458470086]
我々は,拡散モデルのサンプリング過程を高速化するために,確率フロー微分方程式の効率的な解法であるニューラル演算子を用いる。
シーケンシャルな性質を持つ他の高速サンプリング手法と比較して、並列復号法を最初に提案する。
本稿では,CIFAR-10では3.78、ImageNet-64では7.83の最先端FIDを1モデル評価環境で達成することを示す。
論文 参考訳(メタデータ) (2022-11-24T07:30:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。