論文の概要: Optimizers for Diffusion Models: A Controlled Benchmark
- arxiv url: http://arxiv.org/abs/2609.23055v1
- Date: Sat, 19 Sep 2026 14:41:28 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-22 20:29:00.664619
- Title: Optimizers for Diffusion Models: A Controlled Benchmark
- Title(参考訳): 拡散モデルの最適化:制御ベンチマーク
- Abstract要約: 4つの拡散定式化にまたがる制御されたベンチマークを示す。
全ての勝者は3つの種で全予算で再訓練される。
Muon、MARS-M、SOAPはそれぞれ、少なくとも1つの拡散定式化で調整されたAdamWを破った。
- 参考スコア(独自算出の注目度): 75.49404880689772
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://github.com/armanbolatov/diffusion-baselines.
- Abstract(参考訳): 離散拡散モデルは、いくつかのベンチマークで自己回帰言語モデルと一致しているが、それらのトレーニングの最良の方法に関する疑問は、はるかに少ない注目を集めている:最適化器は、ある論文から次の論文に継承され、決して比較されない。
一方、新しいオプティマイザは、損失面の異なる目標である自己回帰事前訓練にほぼ限定して検証される。
7つのオプティマイザ (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) をマスク拡散 (text8), 均一拡散 (QM9, LM1B) と画像上のガウス拡散 (CelebA-64) で表す。
すべてのオプティマイザは、同じ検索プロトコルを受け取り、すべての勝者は、3つのシードで全予算で再訓練される。
AdamWは強力なデフォルトだが、必ずしも正しい選択ではない。それは4つのタスクのうち2つで解決されたマージンに打ち負かされ、勝者は定式化によって変化するため、オプティマイザはトレーニングレシピの他の部分と同じ注意に値する。
Muon, MARS-M, SOAPはそれぞれ,少なくとも1つの拡散定式化で調整されたAdamWを上回った。
ベンチマーク、すべての実行、すべてのフィギュアは、https://github.com/armanbolatov/diffusion-baselines.comでリリースされたコードから終わりまで再現可能である。
関連論文リスト
- Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching [66.39914384073145]
本稿では,安価な拡散サンプリング推論をステップレベル候補の再利用プールに変換する自己整合性フレームワークを提案する。
ステップレベルの再結合は、難しい問題に対して最も有益であることがわかった。
トレーニング不要のフレームワークは、6つの数学およびコーディングタスクの平均精度を最大2倍改善します。
論文 参考訳(メタデータ) (2026-02-26T11:08:39Z) - The Diffusion Duality, Chapter II: $Ψ$-Samplers and Efficient Curriculum [13.49715655470027]
離散拡散のためのプレデクター・コレクター・サンプルのファミリーを紹介する。
均一状態拡散と組み合わせた場合、サンプルは言語と画像のモデリングの両方において祖先サンプリングより優れている。
これらの結果は,Masked 拡散が拡散に基づく言語モデリングの必然的未来であるという仮定を疑問視している。
論文 参考訳(メタデータ) (2026-02-24T18:35:22Z) - Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models [58.946955321428845]
本研究は自己回帰型モンテカルロ(SMC)を提示する。
提案アルゴリズムは,既存のMDLMのほとんどが信頼性に基づくサンプリング戦略に依存している点に起因している。
粒子重み付けのための自己回帰信号として軌道レベルの信頼性を導入する。
論文 参考訳(メタデータ) (2026-02-02T09:21:45Z) - The Diffusion Duality [24.39272541108744]
一様状態拡散過程は、基礎となるガウス拡散から自然に現れる。
カリキュラム学習で訓練されたモデルは、7つのベンチマークのうち3つでゼロショットパープレキシティで自己回帰モデルを上回る。
本稿では, 連続から離散的な状態への連続蒸留を適応させる離散一致蒸留について述べる。
論文 参考訳(メタデータ) (2025-06-12T16:55:35Z) - MARS: Unleashing the Power of Variance Reduction for Training Large Models [56.67982828148859]
深層ニューラルネットワークのための統合トレーニングフレームワークを提案する。
我々は,事前条件付き勾配最適化を利用するMARSの3つの例を紹介する。
その結果,MARSの実装はAdamより一貫して優れていた。
論文 参考訳(メタデータ) (2024-11-15T18:57:39Z) - Variance-reduced Language Pretraining via a Mask Proposal Network [5.819397109258169]
自己指導型学習(英: self-supervised learning, a.k.a.)は、自然言語処理において重要である。
本稿では,勾配分散低減の観点から問題に取り組む。
そこで我々は,マスク提案の最適分布を近似したMAsk Network(MAPNet)を導入した。
論文 参考訳(メタデータ) (2020-08-12T14:12:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。