論文の概要: Avoiding unsafe sets when training with Langevin Dynamics
- arxiv url: http://arxiv.org/abs/2607.07538v2
- Date: Wed, 15 Jul 2026 17:21:26 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-16 14:31:40.665946
- Title: Avoiding unsafe sets when training with Langevin Dynamics
- Title(参考訳): Langevin Dynamics を用いたトレーニングにおける安全でないセットの回避
- Authors: Adam M. Oberman,
- Abstract要約: 形状自由境界の $_t(mathcalA_H) le (mathcalA_H) le (mathcalA_H) le (mathcalA_H) le (mathcalA_H) le (mathcalA_H) le (mathcalA_H) le (mathcalA_H) について検討した。
- 参考スコア(独自算出の注目度): 1.0762008415887194
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics, and a natural safety question is to bound the probability $ν_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in \mathcal{A}_H)$ that the trajectory lies in a designated failure region $\mathcal{A}_H$. We study this for a smooth, strongly convex loss in $d$ dimensions, with $\mathcal{A}_H$ separated from the minimizer by an energy gap. At the end of training, the equilibrium mass $π(\mathcal{A}_H)$ is exponentially small in $d$, with a complementary energy-barrier rate when the noise is small. Along the trajectory, a shape-free bound $ν_t(\mathcal{A}_H) \le π(\mathcal{A}_H)(1 + \sqrt{χ_0^2/π(\mathcal{A}_H)}\,e^{-mt})$ shows the in-set probability relaxes to (twice) the static value after a burn-in of order $d$, using only the global spectral gap $m$. A worked Ornstein-Uhlenbeck example shows this burn-in is necessary: an angular slice of the equilibrium shell can transiently swell by a factor exponential in $d$, though its equilibrium mass is tiny. To rule this out we introduce a local relaxation rate, defined through the spectral measure of the region's centered indicator rather than a Dirichlet-form Rayleigh quotient. For geometrically isolated regions this rate exceeds the global one, shrinking the burn-in, and with a maximum-principle ceiling it caps the trajectory probability uniformly in time. Strong convexity sets how fast training relaxes, but the shape of the unsafe set decides whether the trajectory bulges through it on the way to equilibrium.
- Abstract(参考訳): ノイズ勾配勾配のモデルを訓練することはランゲヴィン力学(英語版)として理想化することができ、自然な安全問題は確率 $ν_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in \mathcal{A}_H)$ を有界化することである。
我々はこれを、エネルギーギャップによって最小値から$\mathcal{A}_H$を分離して、$d$次元の滑らかで強い凸損失を求める。
トレーニングの終わりに、平衡質量$π(\mathcal{A}_H)$は指数関数的に$d$で小さくなり、ノイズが小さいときに相補的なエネルギーバリアレートとなる。
軌道に沿って、形のない有界な$ν_t(\mathcal{A}_H) \le π(\mathcal{A}_H)(1 + \sqrt{*_0^2/π(\mathcal{A}_H)}\,e^{-mt})$ は、大域的スペクトルギャップ$m$のみを用いて、位数$d$のバーンインの後の静的値を(2回)緩和する。
作用したオルンシュタイン=ウレンベックの例では、平衡殻の角のスライスは、平衡質量が小さいにもかかわらず、$d$で指数関数的に過渡的に膨らむことができる。
これを決定するために、ディリクレ形式のレイリー商ではなく、領域の中心的指標のスペクトル測度によって定義される局所緩和率を導入する。
幾何学的に孤立した領域では、この速度はグローバルな領域を超え、バーンインを小さくし、最大主天井で軌道の確率を一定に抑える。
強い凸性は高速なトレーニングの緩和を規定するが、安全でない集合の形状は軌道が平衡に向かう途中で膨らむかどうかを決定する。
関連論文リスト
- A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space [0.0]
ワッサーシュタイン空間において、曲線 $tmapsto_t$ の将来の値 $t_n+h$ を推定するミニマックス速度について検討する。
我々の中心的な結果は、時間空間的ミニマックスの下位境界であり、正規で局所的な輸送に富むサブクラスである。
論文 参考訳(メタデータ) (2026-06-05T14:43:10Z) - Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle [51.714334316332476]
Local Lは制約付き最適化のための新しいプロジェクションフリー型である。
局所LMOはGD(Gradient Descent)のオラクルと見なされる。
論文 参考訳(メタデータ) (2026-05-09T10:03:24Z) - When Does $\ell_2$-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the $\ell_1$ Implicit Bias [15.113649527486276]
良性オーバーフィッティングが線形レートで失敗することを示します。
この局所化機構は信号の存在下で持続するべきであるが、正確な信号-雑音分解は未解決の問題である。
論文 参考訳(メタデータ) (2026-05-07T14:14:09Z) - On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature [1.6773271875801752]
グラディエントDescent (SGD) は、損失ランドスケープの局所曲率と相関する異方性雑音を導入し、平坦なミニマに対して最適化を行う。
この仮定は、ディープニューラルネットワークでは通常違反される制約条件下でのみ成立することを示す。
データセット、アーキテクチャ、損失関数にわたる実験は、これらの境界を検証し、ディープラーニングにおけるノイズ-曲率関係を統一的に評価する。
論文 参考訳(メタデータ) (2026-02-05T12:35:13Z) - Anomalous energy correlations and spectral form factor in the nonergodic phase of the $β$-ensemble [0.0]
臨界エネルギーが特性時間スケールを制御する$beta$-ensembleの力学特性について検討する。
時間窓$t_mathrmR t t_mathrmH$にエネルギー相関が存在しないことを示す。
論文 参考訳(メタデータ) (2025-06-18T09:05:22Z) - Statistical Learning under Heterogeneous Distribution Shift [71.8393170225794]
ground-truth predictor is additive $mathbbE[mathbfz mid mathbfx,mathbfy] = f_star(mathbfx) +g_star(mathbfy)$.
論文 参考訳(メタデータ) (2023-02-27T16:34:21Z) - Random quantum circuits transform local noise into global white noise [118.18170052022323]
低忠実度状態におけるノイズランダム量子回路の測定結果の分布について検討する。
十分に弱くユニタリな局所雑音に対して、一般的なノイズ回路インスタンスの出力分布$p_textnoisy$間の相関(線形クロスエントロピーベンチマークで測定)は指数関数的に減少する。
ノイズが不整合であれば、出力分布は、正確に同じ速度で均一分布の$p_textunif$に近づく。
論文 参考訳(メタデータ) (2021-11-29T19:26:28Z) - Agnostic Learning of Halfspaces with Gradient Descent via Soft Margins [92.7662890047311]
勾配降下は、分類誤差$tilde O(mathsfOPT1/2) + varepsilon$ in $mathrmpoly(d,1/varepsilon)$ time and sample complexity.
論文 参考訳(メタデータ) (2020-10-01T16:48:33Z) - Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and
Variance Reduction [63.41789556777387]
非同期Q-ラーニングはマルコフ決定過程(MDP)の最適行動値関数(またはQ-関数)を学習することを目的としている。
Q-関数の入出力$varepsilon$-正確な推定に必要なサンプルの数は、少なくとも$frac1mu_min (1-gamma)5varepsilon2+ fract_mixmu_min (1-gamma)$の順である。
論文 参考訳(メタデータ) (2020-06-04T17:51:00Z) - Agnostic Learning of a Single Neuron with Gradient Descent [92.7662890047311]
期待される正方形損失から、最も適合した単一ニューロンを学習することの問題点を考察する。
ReLUアクティベーションでは、我々の人口リスク保証は$O(mathsfOPT1/2)+epsilon$である。
ReLUアクティベーションでは、我々の人口リスク保証は$O(mathsfOPT1/2)+epsilon$である。
論文 参考訳(メタデータ) (2020-05-29T07:20:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。