論文の概要: Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics
- arxiv url: http://arxiv.org/abs/2608.04382v1
- Date: Wed, 05 Aug 2026 02:39:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.694413
- Title: Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics
- Title(参考訳): 初期勾配勾配ダイナミクスにおけるロジスティック回帰の非漸近的暗黙バイアス
- Authors: Han Bao,
- Abstract要約: 機械学習から生じる暗黙のバイアスは、しばしば過度なパターンに収まるのを防ぐ。
この研究は、この初期段階アライメント現象のメカニズムを理解することを目的としている。
- 参考スコア(独自算出の注目度): 3.7628827026898466
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization. Implicit bias emerging from optimization, though not being encoded by the learning objective, often prevents from overfitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of "train longer, generalize better." However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than pure convex optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within $O(\exp(\exp(-δ)))$ iterations, where $δ>0$ is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to establishing faster weak alignment.
- Abstract(参考訳): 勾配降下は、最適化にのみ焦点をあてるだけでなく、現代の機械学習において特に関心を寄せている。
学習目的によって符号化されないにもかかわらず、最適化から生じる暗黙のバイアスは、しばしば過度に適合し、刺激的なパターンを妨げます。
典型的な例は、指数関数的に尾行された損失関数に対して広く確立された線形分類器の極大陰性バイアスである。
与えられたデータセットを分離した後でも、パラメータベクトルは勾配降下ダイナミクスに沿って漸近的に最大辺方向に向かって進化し続けている。
この現象は、しばしば「より長く、より一般化する」という経験的な観察を裏付けるものである。
しかし、最大マージン収束は漸近現象であり、さらに悪いことに、この漸近収束速度は純粋な凸最適化よりも著しく遅い。
それでも、勾配勾配勾配のダイナミックスに沿ったパラメータベクトルは、漸近速度よりもかなり少ないイテレーションで(正確にはそうではないが)正のマルジン方向と相関するのが一般的である。
この古典的な問題に別の光を当てることで、この初期のアライメント現象のメカニズムを理解することを目的としている。
我々の理論的な結果は、パラメータベクトルが$O(\exp(\exp(-δ))$ iterations 内で最大辺方向と弱整合していることを示し、$δ>0$ は許容アライメント誤差であり、これは厳密であることが示されている。
放射と接する流れを追跡することで、我々の証明はデータセットの幾何と直接的にアライメントのダイナミクスを実行し、漸近的拡張を排除し、より高速な弱いアライメントを確立するための重要な洞察となる。
関連論文リスト
- High-dimensional Limit of SGD for Diagonal Linear Networks [20.199898128645497]
対角線ネットワーク上の勾配勾配は微分方程式(SDE)によって支配される連続力学によりよく近似されることを示す。
適切なパラメトリゼーションの下では、このダイナミクスは地球規模で十分に仮定され、高い確率で指数関数的に高速にゼロリスクに収束し、それらの長時間の振る舞いを完全に明示的な非漸近的記述をもたらすことが示される。
論文 参考訳(メタデータ) (2026-05-16T22:26:59Z) - On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes [15.63629978994481]
我々はこの勾配カスケード設定に関する最初の包括的な理論的解析を提示する。
摂動が勾配収束順序を悪化させない条件を特定する。
論文 参考訳(メタデータ) (2026-02-24T07:47:15Z) - Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations [57.179679246370114]
乱摂動の分布は, 摂動段差がゼロになる傾向にあるため, 推定子の分散を最小限に抑える。
以上の結果から, 一定の長さを維持するのではなく, 真の勾配に方向を合わせることが可能であることが示唆された。
論文 参考訳(メタデータ) (2025-10-22T19:06:39Z) - Long-time dynamics and universality of nonconvex gradient descent [0.7614628596146601]
本稿では,非勾配勾配の長期的挙動を単一インデックスモデルで特徴づけるための一般的な手法を開発する。
我々のアプローチでは、勾配降下は概してデータとは独立であり、特徴ベクトルと強く一致しないことが明らかとなった。
論文 参考訳(メタデータ) (2025-09-14T20:36:18Z) - Implicit Bias and Fast Convergence Rates for Self-attention [26.766649949420746]
本稿では,変圧器の定義機構である自己注意の基本的な最適化原理について考察する。
線形分類におけるデコーダを用いた自己アテンション層における勾配ベースの暗黙バイアスを解析する。
論文 参考訳(メタデータ) (2024-02-08T15:15:09Z) - Good regularity creates large learning rate implicit biases: edge of
stability, balancing, and catapult [49.8719617899285]
非最適化のための客観的降下に適用された大きな学習速度は、安定性の端を含む様々な暗黙のバイアスをもたらす。
この論文は降下の初期段階を示し、これらの暗黙の偏見が実際には同じ氷山であることを示す。
論文 参考訳(メタデータ) (2023-10-26T01:11:17Z) - Implicit Bias of Gradient Descent for Logistic Regression at the Edge of
Stability [69.01076284478151]
機械学習の最適化において、勾配降下(GD)はしばしば安定性の端(EoS)で動く
本稿では,EoS系における線形分離可能なデータに対するロジスティック回帰のための定数段差GDの収束と暗黙バイアスについて検討する。
論文 参考訳(メタデータ) (2023-05-19T16:24:47Z) - Proximal Subgradient Norm Minimization of ISTA and FISTA [8.261388753972234]
反復収縮保持アルゴリズムのクラスに対する2乗近位次数ノルムは逆2乗率で収束することを示す。
また、高速反復収縮保持アルゴリズム (FISTA) のクラスに対する2乗次次数次ノルムが、逆立方レートで収束するように加速されることも示している。
論文 参考訳(メタデータ) (2022-11-03T06:50:19Z) - Nonconvex Stochastic Scaled-Gradient Descent and Generalized Eigenvector
Problems [98.34292831923335]
オンライン相関解析の問題から,emphStochastic Scaled-Gradient Descent (SSD)アルゴリズムを提案する。
我々はこれらのアイデアをオンライン相関解析に適用し、局所収束率を正規性に比例した最適な1時間スケールのアルゴリズムを初めて導いた。
論文 参考訳(メタデータ) (2021-12-29T18:46:52Z) - Direction Matters: On the Implicit Bias of Stochastic Gradient Descent
with Moderate Learning Rate [105.62979485062756]
本稿では,中等度学習におけるSGDの特定の正規化効果を特徴付けることを試みる。
SGDはデータ行列の大きな固有値方向に沿って収束し、GDは小さな固有値方向に沿って収束することを示す。
論文 参考訳(メタデータ) (2020-11-04T21:07:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。