論文の概要: The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks
- arxiv url: http://arxiv.org/abs/2610.09132v1
- Date: Tue, 06 Oct 2026 21:24:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 21:58:22.611766
- Title: The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks
- Title(参考訳): 変分ベイズ線形ニューラルネットワークの極限予測モーメントに及ぼす等温性の影響
- Abstract要約: ワイドベイズニューラルネットワークでは、ガウス平均場変動推論は「優越性」の傾向にある
我々は、$T = /Mc$, with constants $, c > 0$, as $M to infty$という形式のスケジュール下での限定予測分布を導出し、それを正確な後部の無限幅極限である未干渉ニューラルネットワークガウス過程(NNGP)の後方と比較する。
- 参考スコア(独自算出の注目度): 2.37685358666088
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to "prior dominance": the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width $M$ grows. Tempering the likelihood, by raising it to the power $1/T$ for a temperature $T < 1$, is equivalent to scaling the KL term by $T$. We ask in this paper how fast $T$ must decrease with $M$ to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form $T = τ/M^{c}$, with constants $τ, c > 0$, as $M \to \infty$ and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at $c = 1/2$, once $τ$ falls below an explicit threshold, and equals the least-squares prediction for $c > 1/2$, whereas the limiting variance keeps its prior value for $c < 1$, matches the NNGP's for $c=1$, and vanishes for $c > 1$. With suitable choices of $τ,c$, one can recover either the NNGP posterior expectation or its variance.
- Abstract(参考訳): 広義ベイズニューラルネットワークでは、ガウス平均場変量推論が「優位性」に陥りやすい: ELBOのクルバック・リーバー(KL)正則化項は、期待される対数類似度を上回っ、その変動予測分布は、幅$M$が大きくなるにつれて、先行予測値に崩壊する。
温度$T < 1$で1/T$に上げると、KL項を$T$にスケールする。
この論文では、この縮退に対抗するために、$T$が$M$でどれだけ早く減少し、この2つの項の間に良いバランスをとらなければならないか尋ねる。
等方性ガウス事前を持つ単層線形ネットワークの場合、定数$τ, c > 0$ を $M \to \infty$ とする$T = τ/M^{c}$ という形のスケジュールの下で予測分布の制限を導出し、それを未測定のニューラルネットワークガウス過程(NNGP)後部と比較する。
制限期待値が$c = 1/2$、一度$τ$が明示的なしきい値を下回り、最小二乗予想値が$c > 1/2$と等しいのに対して、制限偏差値が$c < 1$、NNGPの$c=1$、および$c > 1$と一致する。
適切な$τ,c$を選択すると、NNGP後続期待値またはその分散値を回復することができる。
関連論文リスト
- From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting [7.309983740889802]
本研究では,従来の定常予測理論により,よりコンパクトな線形予測器のパラメータ共有を導出する方法について検討する。
我々のHankel-Toeplitz Forecasterは、逆フィルタと予測マップの両方を定義するインパルス応答を学習する。
論文 参考訳(メタデータ) (2026-09-27T22:35:52Z) - Constrained Online Learning with Noisy Constraint Values [55.29259818039367]
一般的な実現可能性の下では、我々のLEDGERアルゴリズムは、期待される損失$O(sqrt T)と期待される予算違反$O(sqrtTlog(eT))を達成します。
スレーター条件、フィードバックチャネル間の独立性、絶対的制約値境界は不要である。
論文 参考訳(メタデータ) (2026-09-07T01:38:41Z) - How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis [53.063298916923976]
r*(x) = *(langle *, xrangle)$ と $x sim N(0, I_d)$ でガウスの単一インデックスモデルでフィードバックを研究する。
まず、報酬重み付きサンプルから隠れた方向を*$で学習し、次に重み付きリッジ回帰により読み出し層に適合する2段階のニューラル報酬モデルを分析する。
論文 参考訳(メタデータ) (2026-05-23T22:00:38Z) - When Does $\ell_2$-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the $\ell_1$ Implicit Bias [15.113649527486276]
良性オーバーフィッティングが線形レートで失敗することを示します。
この局所化機構は信号の存在下で持続するべきであるが、正確な信号-雑音分解は未解決の問題である。
論文 参考訳(メタデータ) (2026-05-07T14:14:09Z) - Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification [50.717692060500696]
対数損失を伴う次のトーケン予測は自己回帰シーケンスモデリングの基盤となる。
次トーケン予測は、適度な誤差増幅を表す$C=tilde O(H)$を達成するために堅牢にすることができる。
C=e(log H)1-Omega(1)$。
論文 参考訳(メタデータ) (2025-02-18T02:52:00Z) - Mind the Gap: A Causal Perspective on Bias Amplification in Prediction & Decision-Making [58.06306331390586]
本稿では,閾値演算による予測値がS$変化の程度を測るマージン補数の概念を導入する。
適切な因果仮定の下では、予測スコア$S$に対する$X$の影響は、真の結果$Y$に対する$X$の影響に等しいことを示す。
論文 参考訳(メタデータ) (2024-05-24T11:22:19Z) - $L^1$ Estimation: On the Optimality of Linear Estimators [64.76492306585168]
この研究は、条件中央値の線型性を誘導する$X$上の唯一の先行分布がガウス分布であることを示している。
特に、条件分布 $P_X|Y=y$ がすべての$y$に対して対称であるなら、$X$ はガウス分布に従う必要がある。
論文 参考訳(メタデータ) (2023-09-17T01:45:13Z) - Exact one- and two-site reduced dynamics in a finite-size quantum Ising
ring after a quench: A semi-analytical approach [4.911435444514558]
クエンチ後の等質量子イジング環の非平衡ダイナミクスについて検討する。
1つのスピンと2つの最も近い隣り合うスピンの長時間還元ダイナミクスについて研究した。
論文 参考訳(メタデータ) (2021-03-23T13:14:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。