論文の概要: Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum
- arxiv url: http://arxiv.org/abs/2609.39595v1
- Date: Wed, 30 Sep 2026 12:19:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:27.687128
- Title: Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum
- Title(参考訳): 有限Newton-Schulz反復とNesterov Momentumによる実用ムーンの収束
- Abstract要約: 実践的なミューオンは運動量を維持し、各パラメータ行列に対して少数のニュートン-シュルツ反復を実行する。
- 参考スコア(独自算出の注目度): 11.807444451741851
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.
- Abstract(参考訳): 実践的なミューオンは運動量を維持し、各パラメータ行列に対して小さな一定数のニュートン-シュルツ反復を実行する。
我々はこれらの階層的な有限ステップ更新を、正確な極性因子や1つの大域直交化によって置き換えるのではなく、結合された非凸目的に基づいて共同で解析する。
勾配依存$(\mathcal L_0,\mathcal L_1,q)$-smoothness および有界層次分散を持つ条件付き非バイアス確率勾配の下では、期待平均フロベニウス勾配ノルムに有界な$\mathcal O(T^{-1/4})$を確立する。
この解析はネステロフ再帰を保ち、非ゼロ出力特異値上の有界確率勾配、対称雑音、および一様正下界を必要としない。
その定数は、ブロックの数と問題定数が固定されたとき、明示的な行列次元やランク要素を含まない。
この証明は、降下不等式と運動量追跡誤差の初期化、ノイズ、ドリフトへの分解に続く。
元の5段階のクインティックに対して、必要なスカラーマップ境界を解析的に検証し、同じ境界を満たすステップ依存係数も許容する。
相補的な核ノルムの結果は、強いスペクトル条件下でランク依存を定量化する。
消滅率には、標準の単一係数ネステロフ規則を含む学習速度と運動量スケジュールが混在している。
関連論文リスト
- Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks [16.84694273405234]
ヒンジロスを伴う2層2層活性化ネットワークを訓練するための恒常的ストレートスルー推定器(STE)について検討した。
我々の中心的な問題は、アルゴリズム安定性が不連続なSTEトレーニングルールによって生成される推定器の統計的一般化を説明することができるかどうかである。
論文 参考訳(メタデータ) (2026-09-06T07:12:16Z) - A proximal subgradient method for nonconvex stochastic optimization under the Kurdyka-Łojasiewicz condition [39.146761527401424]
本研究は,期待されるコストの和を最小化するための近位次法を導入する。
クルディカ・ロジャシエヴィチ(KL)特性を用いて、全軌道の1つの定常点への収束を保証する。
論文 参考訳(メタデータ) (2026-08-05T23:07:27Z) - Discretization and Statistical Consistency of Functional Flow Matching [0.0]
有限ランク再構成の強い一貫した列に対して、有限条件速度目標の強い$L2$収束を証明した。
学習フローに対して、集団重ね合わせ経路に直接結合すると、終端ワッサーシュタイン境界が得られる。
論文 参考訳(メタデータ) (2026-08-05T07:05:47Z) - Variance Reduction for Stochastic Gradient Generalized Non-reversible Langevin Monte Carlo Algorithms [7.336637417508747]
非可逆ランゲヴィン力学に対する勾配丸山推定器の序列変動について検討する。
我々は、逆二乗階段の水平方向上の経験的平均が、消滅段数状態における中心極限定理を満たすことを証明した。
論文 参考訳(メタデータ) (2026-06-27T08:25:48Z) - A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling [7.035974899001363]
我々は, 隣接再帰安定性と重み付き濃度ステップから, 有限ラウンド上尾保証を導出した。
一次元の反例は、なぜギャップ、平滑化、あるいは正則性条件が必要なのかを示す。
論文 参考訳(メタデータ) (2026-06-01T05:36:26Z) - High-Probability Bounds for Stochastic Optimization and Variational
Inequalities: the Case of Unbounded Variance [59.211456992422136]
制約の少ない仮定の下で高確率収束結果のアルゴリズムを提案する。
これらの結果は、標準機能クラスに適合しない問題を最適化するために検討された手法の使用を正当化する。
論文 参考訳(メタデータ) (2023-02-02T10:37:23Z) - Optimal policy evaluation using kernel-based temporal difference methods [78.83926562536791]
カーネルヒルベルト空間を用いて、無限水平割引マルコフ報酬過程の値関数を推定する。
我々は、関連するカーネル演算子の固有値に明示的に依存した誤差の非漸近上界を導出する。
MRP のサブクラスに対する minimax の下位境界を証明する。
論文 参考訳(メタデータ) (2021-09-24T14:48:20Z) - Spectral clustering under degree heterogeneity: a case for the random
walk Laplacian [83.79286663107845]
本稿では,ランダムウォークラプラシアンを用いたグラフスペクトル埋め込みが,ノード次数に対して完全に補正されたベクトル表現を生成することを示す。
次数補正ブロックモデルの特別な場合、埋め込みはK個の異なる点に集中し、コミュニティを表す。
論文 参考訳(メタデータ) (2021-05-03T16:36:27Z) - On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and
Non-Asymptotic Concentration [115.1954841020189]
The inequality and non-asymptotic properties of approximation procedure with Polyak-Ruppert averaging。
一定のステップサイズと無限大となる反復数を持つ平均的反復数に対する中心極限定理(CLT)を証明する。
論文 参考訳(メタデータ) (2020-04-09T17:54:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。