論文の概要: Correlated initialization of deep residual networks
- arxiv url: http://arxiv.org/abs/2609.03589v1
- Date: Thu, 03 Sep 2026 09:35:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-04 18:28:39.004178
- Title: Correlated initialization of deep residual networks
- Title(参考訳): 深部残留ネットワークの相関初期化
- Abstract要約: 初期化時に層間に重みが相関する残余ネットワークの大規模挙動について検討した。
我々は、無限深度極限がエルミート過程によって駆動されるヤング微分方程式の解であるようなユニークな臨界スケーリングが存在することを証明した。
我々は、バナッハ空間におけるヤング微分方程式の頑健な安定性理論を確立する新しい結果の集合に依存している。
- 参考スコア(独自算出の注目度): 1.9116784879310027
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.
- Abstract(参考訳): 初期化時に層間に重みが相関する残余ネットワークの大規模挙動について検討した。
この結果から, 独立初期化から生じるブラウン確率微分方程式と完全相関初期化から生じる常微分方程式との間に, 相関初期化が連続的に介在するマリオン等 [2025] の予想を確認し, 拡張する。
特徴関数の定常ガウス列への定常的相関による初期化が得られれば、無限深度極限がエルミート過程によって駆動されるヤング微分方程式の解となるようなユニークな臨界スケーリングが存在することが証明される。
エルミート過程は、初期化を生成する特徴関数がエルミート階数 1 を持つならば、分数的ブラウン運動に還元される。
臨界スケーリングと漸近極限は、特徴関数のエルミート階数とともに相関の減衰によって一意に決定されることを示す。
その結果、初期化の相関構造とエルミートランクは、漸近的状態における有意義なハイパーパラメータを表す。
対照的に、有限分散 iid 初期化の下では、漸近ドライバは分布の選択にかかわらず、普遍的にブラウン化である。
我々の証明はバナッハ空間におけるヤング微分方程式の頑健な安定性理論を確立する新しい結果の集合に依存する。
関連論文リスト
- Quotient Dynamics, Effective Curvature, and Implicit Bias in Positive Quadratic Networks [0.0]
正の二次ネットワークは、低ランク表現 f_U(x)=xtop UUtop x を持ち、UinmathbbRdtimes r は右乗法までしか特定できない。
この商構造は, トレーニング力学, 曲率, 回復, バイアスをいかに支配するかを考察する。
論文 参考訳(メタデータ) (2026-07-28T12:04:46Z) - Generalized nonparametric regression in reproducing kernel Hilbert spaces: Consistency and rates of convergence [0.0]
我々は再生空間におけるM-levの包括的理論を開発する。
ソボ空間に対しては、混合滑らか性関数を持つ滑らか性空間に接続する新たなレートを得る。
論文 参考訳(メタデータ) (2026-06-22T08:12:31Z) - Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum [6.749750044497731]
我々は、ガウス以下の設計仮定の下で、誘惑的な過フィットと破滅的な過フィットの現象を証明した。
また、機能の独立性は、誘惑に満ちたオーバーフィッティングを保証する上で重要な役割を担っていることも確認しています。
論文 参考訳(メタデータ) (2024-02-02T10:36:53Z) - On the Identification and Optimization of Nonsmooth Superposition
Operators in Semilinear Elliptic PDEs [3.045851438458641]
原型半線形楕円偏微分方程式(PDE)の非線形部分におけるネミトスキー作用素の同定を目的とした無限次元最適化問題について検討する。
以前の研究とは対照的に、ネミトスキー作用素を誘導する関数が a-priori であることは、$H leakyloc(mathbbR)$ の要素であることが知られている。
論文 参考訳(メタデータ) (2023-06-08T13:33:20Z) - Kernel-based off-policy estimation without overlap: Instance optimality
beyond semiparametric efficiency [53.90687548731265]
本研究では,観測データに基づいて線形関数を推定するための最適手順について検討する。
任意の凸および対称函数クラス $mathcalF$ に対して、平均二乗誤差で有界な非漸近局所ミニマックスを導出する。
論文 参考訳(メタデータ) (2023-01-16T02:57:37Z) - Variational Representations of Annealing Paths: Bregman Information
under Monotonic Embedding [12.020235141059992]
算術平均は、予想されるブレグマン偏差を1つの代表点に最小化することを示す。
本分析では, 準算術的手段, パラメトリック・ファミリー, 発散関数の相互作用に注目した。
論文 参考訳(メタデータ) (2022-09-15T17:22:04Z) - Nonconvex Stochastic Scaled-Gradient Descent and Generalized Eigenvector
Problems [98.34292831923335]
オンライン相関解析の問題から,emphStochastic Scaled-Gradient Descent (SSD)アルゴリズムを提案する。
我々はこれらのアイデアをオンライン相関解析に適用し、局所収束率を正規性に比例した最適な1時間スケールのアルゴリズムを初めて導いた。
論文 参考訳(メタデータ) (2021-12-29T18:46:52Z) - Spectral clustering under degree heterogeneity: a case for the random
walk Laplacian [83.79286663107845]
本稿では,ランダムウォークラプラシアンを用いたグラフスペクトル埋め込みが,ノード次数に対して完全に補正されたベクトル表現を生成することを示す。
次数補正ブロックモデルの特別な場合、埋め込みはK個の異なる点に集中し、コミュニティを表す。
論文 参考訳(メタデータ) (2021-05-03T16:36:27Z) - Determination of the critical exponents in dissipative phase
transitions: Coherent anomaly approach [51.819912248960804]
オープン量子多体系の定常状態に存在する相転移の臨界指数を抽出するコヒーレント異常法の一般化を提案する。
論文 参考訳(メタデータ) (2021-03-12T13:16:18Z) - On the Implicit Bias of Initialization Shape: Beyond Infinitesimal
Mirror Descent [55.96478231566129]
学習モデルを決定する上で,相対スケールが重要な役割を果たすことを示す。
勾配流の誘導バイアスを導出する手法を開発した。
論文 参考訳(メタデータ) (2021-02-19T07:10:48Z) - Approximation Schemes for ReLU Regression [80.33702497406632]
我々はReLU回帰の根本的な問題を考察する。
目的は、未知の分布から引き出された2乗損失に対して、最も適したReLUを出力することである。
論文 参考訳(メタデータ) (2020-05-26T16:26:17Z) - On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and
Non-Asymptotic Concentration [115.1954841020189]
The inequality and non-asymptotic properties of approximation procedure with Polyak-Ruppert averaging。
一定のステップサイズと無限大となる反復数を持つ平均的反復数に対する中心極限定理(CLT)を証明する。
論文 参考訳(メタデータ) (2020-04-09T17:54:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。