論文の概要: Scaling Limits of Constant-Stepsize SGD at Flat Minima
- arxiv url: http://arxiv.org/abs/2607.16384v1
- Date: Fri, 17 Jul 2026 17:22:46 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.138217
- Title: Scaling Limits of Constant-Stepsize SGD at Flat Minima
- Title(参考訳): 定段SGDのフラットミニマにおけるスケーリング限界
- Abstract要約: 契約駆動チェーンが生成するマルコフ雑音について検討する。
定段数$$の勾配降下(SGD)に対して、最小値を中心とする反復法の不変法則は、長い時間的地平線上でのアルゴリズムの振舞いを記述する。
この振舞いは、平坦なミニマと(部分)四角形尾を持つ凸対象に対して根本的に変化することを示す。
- 参考スコア(独自算出の注目度): 5.608222858044445
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: For stochastic gradient descent (SGD) with a constant stepsize $α$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar $\sqrtα$ scaling and a Gaussian limit as $α\downarrow 0$. We show that this behavior changes fundamentally for convex objectives $H$ with flat minima and (sub)quadratic tails. More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize $α$, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an $α$-dependent metric. When the minimizer $x_\star$ has local flatness exponent $m\ge2$, meaning that $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ as $x\to x_\star$, we obtain a contraction bound with factor $1-cα^{m-1}$, where $c>0$ is a constant. This recovers the factor $1-cα$ in the quadratic case $m=2$. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale $α^{1/m}$ and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation $$ dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t , $$ where $h_0$ is the limiting drift at the minimizer and $Σ$ denotes the asymptotic covariance. This recovers the Gaussian limit when $m=2$ and gives generally non-Gaussian stationary limits in the flat case $m>2$. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.
- Abstract(参考訳): 定段数$α$の確率勾配降下(SGD)に対して、最小値を中心とする反復の不変法則は、長い時間的地平線上でのアルゴリズムの振舞いを記述している。
強凸の場合、この不変法則はよく知られた$\sqrtα$スケーリングを持ち、ガウス極限は$α\downarrow 0$である。
この振舞いは、平坦なミニマと(部分)四角形尾を持つ凸対象に対して根本的に変化することを示す。
具体的には,契約駆動チェーンが生成するマルコフ雑音を用いたSGDについて検討する。
十分小さな定数の次数$α$に対して、$α$依存計量によって誘導されるワッサーシュタイン距離における拡張不変法則への存在、一意性、幾何学的収束を証明する。
最小値 $x_\star$ が局所平坦指数 $m\ge2$ を持ち、つまり $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ が $x\to x_\star$ であるなら、$c>0$ は定数である。
これは二次の場合、$m=2$で1-cα$を回復する。
次に、小規模スケーリングの限界を分析する。
不変法則はスケール$α^{1/m}$に集中し、再スケール反復は確率微分方程式$$$dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t ,$$$,$h_0$は最小値の極限ドリフトであり、$Σ$は漸近共分散を表す。
これは、$m=2$のときガウス極限を回復し、フラットケース$m>2$のとき、一般に非ガウス定常極限を与える。
最後に、不等平坦指数を持つ座標分離対象に対して、対応する結果を与える。
関連論文リスト
- One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion [0.0]
1つの暗黙のDDIM反転ステップは、事前訓練された拡散モデルが局所多様体幾何学を符号化するかどうかの最も安価なプローブである。
フェルミ窓はモデルのトレーニングサポートと3.6$-$5.6times$で対立し、ヘッセン=リプシッツ定数は法が読み取る曲率の2-$12%である。
最後の条件付き天井 $_t_max(mathrmsym,J)le1$, from $mathrmCov(x_0mid x_t)succeq0$
論文 参考訳(メタデータ) (2026-08-24T11:01:54Z) - High Minima of Gaussian Processes: Overshoots and Minimizer Locations [0.0]
条件付き$M>u$の場合、スケールされたオーバーシュート$u(M-u)$は、平均$_*2$の指数確率変数に$utoinfty$として収束することを示す。
結果は、定常ガウス過程、分数ブラウン運動、分数ブラウンシートによって説明される。
論文 参考訳(メタデータ) (2026-07-22T20:31:36Z) - Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence [54.59847568544922]
有限水平時間同質なマルコフ決定過程に対して、$A$状態、$A$アクション、hoighty $H$、および1ドルで有界なトラジェクティブ当たりの合計報酬について、地平自由な後悔について検討する。
失敗確率$$K$はエピソード数で$tilde O(sqrtSAK+S3K)$ hides $mathsfpolyである。
論文 参考訳(メタデータ) (2026-07-22T07:42:19Z) - Solving Stochastic Fixed-Point Equations with High Probability [22.376855234542813]
オラクルの不動点方程式 $mathbfT(mathbfx) = mathbfx$ をノルム空間上で研究する。
本稿では,2次スムーズなバナッハ空間に対する分散還元段階Halpern法であるVR-GHALを紹介する。
論文 参考訳(メタデータ) (2026-07-10T04:59:20Z) - Estimation of the sub-Gaussian parameter [7.4076878426925035]
準ガウス確率変数 $X$ は 2_* = s_upin mathbbR L()$ ここで $L() = frac22 log mathbbE eX$ は重み付き累積生成関数である。
ガウス以下の確率変数が多用されているにもかかわらず、$2_*$の推定はほとんど注目されず、まだよく理解されていない。
論文 参考訳(メタデータ) (2026-06-04T16:48:31Z) - When Does $\ell_2$-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the $\ell_1$ Implicit Bias [15.113649527486276]
良性オーバーフィッティングが線形レートで失敗することを示します。
この局所化機構は信号の存在下で持続するべきであるが、正確な信号-雑音分解は未解決の問題である。
論文 参考訳(メタデータ) (2026-05-07T14:14:09Z) - Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy [0.0]
クロスエントロピートレーニング分析は、提案されたステップが目標を減少させるかどうかを予測するために、損失の局所的なテイラーモデルに依存する。
提案した更新方向に沿って,ロジット線形化の下で閉形式式を導出する。
_a$の正規化は、標準偏差$0.992$から$0.164$へのオンセット閾値の広がりを縮小する。
論文 参考訳(メタデータ) (2026-03-13T19:42:12Z) - Near-Optimal Convergence of Accelerated Gradient Methods under Generalized and $(L_0, L_1)$-Smoothness [57.93371273485736]
我々は、最近提案された$ell$-smoothness条件$|nabla2f(x)|| le ellleft(||nabla f(x)||right),$$$L$-smoothnessと$(L_0,L_1)$-smoothnessを一般化する関数を持つ凸最適化問題の一階法について検討する。
論文 参考訳(メタデータ) (2025-08-09T08:28:06Z) - On the $O(\frac{\sqrt{d}}{T^{1/4}})$ Convergence Rate of RMSProp and Its Momentum Extension Measured by $\ell_1$ Norm [54.28350823319057]
本稿では、RMSPropとその運動量拡張を考察し、$frac1Tsum_k=1Tの収束速度を確立する。
我々の収束率は、次元$d$を除くすべての係数に関して下界と一致する。
収束率は$frac1Tsum_k=1Tと類似していると考えられる。
論文 参考訳(メタデータ) (2024-02-01T07:21:32Z) - Measurement-induced phase transition for free fermions above one dimension [46.176861415532095]
自由フェルミオンモデルに対する$d>1$次元における測定誘起エンタングルメント相転移の理論を開発した。
臨界点は、粒子数と絡み合いエントロピーの第2累積のスケーリング$$elld-1 ln ell$でギャップのない位相を分離する。
論文 参考訳(メタデータ) (2023-09-21T18:11:04Z) - A first-order primal-dual method with adaptivity to local smoothness [64.62056765216386]
凸凹対象 $min_x max_y f(x) + langle Ax, yrangle - g*(y)$, ここで、$f$ は局所リプシッツ勾配を持つ凸関数であり、$g$ は凸かつ非滑らかである。
主勾配ステップと2段ステップを交互に交互に行うCondat-Vuアルゴリズムの適応バージョンを提案する。
論文 参考訳(メタデータ) (2021-10-28T14:19:30Z) - Private Stochastic Convex Optimization: Optimal Rates in $\ell_1$
Geometry [69.24618367447101]
対数要因まで $(varepsilon,delta)$-differently private の最適過剰人口損失は $sqrtlog(d)/n + sqrtd/varepsilon n.$ です。
損失関数がさらなる滑らかさの仮定を満たすとき、余剰損失は$sqrtlog(d)/n + (log(d)/varepsilon n)2/3で上界(対数因子まで)であることが示される。
論文 参考訳(メタデータ) (2021-03-02T06:53:44Z) - A Simple Convergence Proof of Adam and Adagrad [74.24716715922759]
我々はAdam Adagradと$O(d(N)/st)$アルゴリズムの収束の証明を示す。
Adamはデフォルトパラメータで使用する場合と同じ収束$O(d(N)/st)$で収束する。
論文 参考訳(メタデータ) (2020-03-05T01:56:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。