論文の概要: Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
- arxiv url: http://arxiv.org/abs/2610.02355v1
- Date: Thu, 01 Oct 2026 18:32:34 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.04499
- Title: Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
- Title(参考訳): 適応バッチはLCMの事前学習になぜ役立つのか : 非有界変数の視点から
- Abstract要約: 10トレーニング中のバッチサイズの増加は、事前トレーニングの一般的なプラクティスである。
本稿では,実用的な騒音の挙動をより厳密に記述するノイズモデルを提案する。
- 参考スコア(独自算出の注目度): 7.259329208302806
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence suggests that this assumption fails in many practical nonconvex problems. The Blum--Gladyshev (BG-$0$) noise model relaxes this assumption by allowing the variance to grow quadratically with the distance from initialization, suggesting that batch size schedulers can help by controlling the variance growth during training. However, this growth can be overly conservative in practice. We empirically investigate variance growth in LLM pretraining and observe that a generalized BG model with a tunable growth exponent provides a tighter description of practical noise behavior. Motivated by this observation, we introduce the generalized BG-$a$ noise model, which interpolates between bounded variance ($a=0$) and BG-$0$ noise ($a=2$). Under $L$-smoothness, we derive an information-theoretic lower bound with growth-dependent oracle complexity $Ω(ε^{-(4+a)})$ and establish a matching upper bound in $ε$-dependence by increasing the batch size as the iterates move away from initialization. Finally, we propose an adaptive batch scheduler that controls variance growth through dynamic batch size adjustments during training. In pretraining OLMo2 models of up to 1B parameters on C4, our scheduler achieves a lower validation loss than both small and large batch training under matched token budgets, while using less than 10\% of the iterations of small batch training.
- Abstract(参考訳): トレーニング中のバッチサイズの増加は、大規模言語モデル(LLM)事前トレーニングにおいて一般的なプラクティスであるが、その成功の理論的根拠はよく理解されていない。
確率最適化の分析は、しばしば一様有界な確率勾配の分散を仮定するが、最近の証拠は、この仮定が多くの実用的な非凸問題で失敗することを示唆している。
Blum-Gladyshev (BG-$0$) ノイズモデルはこの仮定を緩和し、分散を初期化からの距離で二次的に成長させることで、バッチサイズスケジューラがトレーニング中の分散成長を制御するのに役立つことを示唆している。
しかし、この成長は実際には過度に保守的である。
我々は, LLM事前学習における分散成長を実験的に検討し, チューナブル成長指数を持つ一般化BGモデルが, 実用的な騒音挙動のより厳密な記述を提供することを示した。
この観測から得られた一般BG-$a$ノイズモデルを導入し,有界分散(a=0$)とBG-$0$ノイズ(a=2$)を補間する。
L$-smoothnessの下では、成長依存のオラクル複雑性を持つ情報理論の下限を$Ω(ε^{-(4+a)})$で導出し、反復が初期化から離れるにつれてバッチサイズを増大させることで、一致する上限を$ε$-dependenceで確立する。
最後に,適応型バッチスケジューラを提案する。
C4上で最大1BパラメータのOLMo2モデルを事前トレーニングする場合、スケジューラは、小さなバッチトレーニングのイテレーションの10倍未満を使用しながら、マッチしたトークン予算下での小規模および大規模バッチトレーニングよりも低い検証損失を達成する。
関連論文リスト
- Double Descent and Malign Overfitting in Diffusion Models [5.939780039158003]
拡散モデルの過度な適合は破滅的であり、モデルを体制へと駆り立てる。
トレーニングサンプルあたりの雑音実現の固定数$m$では、2次ピークが発生するが、標準回帰のように$psim n$ではなく$psim nm$となる。
この過度な適合は、トレーニングの暗黙の規則化が完全に作業中であるにもかかわらず、真のスコアではなく、経験的なスコアに向かってモデルを駆動するからである。
論文 参考訳(メタデータ) (2026-09-22T13:30:48Z) - REVES: REvision and VErification--Augmented Training for Test-Time Scaling [53.197756110943395]
本稿では,オンラインデータ/プロンプト拡張とポリシー最適化を交互に行う2段階反復フレームワークを提案する。
我々は、RLベースライン上の+6.5点と、標準マルチターントレーニングにおける+4.0点の利得を観察する。
論文 参考訳(メタデータ) (2026-06-17T10:37:23Z) - The Effect of Mini-Batch Noise on the Implicit Bias of Adam [2.8647133890966994]
本稿では,ミニバッチノイズがAdamの暗黙の記憶バイアスにどのように影響するかを理解するための理論的枠組みを提案する。
大規模なバッチサイズの場合、メモリによる反正則化の規模が高くなる(一般化を促す)が、バッチサイズが小さくなると($に対する反正則化)依存が逆になる。
我々の一般化は、臨界バッチサイズのスケールにシフトするバッチサイズのスケールを結びつけます。
論文 参考訳(メタデータ) (2026-02-02T04:59:24Z) - Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling [75.36692892951018]
トレーニング中のバッチサイズの増加は、大規模な言語モデルの事前トレーニングを加速するための有望な戦略である。
この研究はバッチサイズスケジューリングのための原則化されたフレームワークを開発する。
標準スケジューラが学習率を半減するたびに、Seesawは1/sqrt2$と倍増し、バッチサイズを倍増します。
論文 参考訳(メタデータ) (2025-10-16T14:17:38Z) - Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model [89.8764435351222]
分散を低減した行列生成のために, WTA-CRS と呼ばれる新しい非バイアス推定系を提案する。
我々の研究は、チューニング変換器の文脈において、提案した推定器が既存のものよりも低い分散を示すという理論的および実験的証拠を提供する。
論文 参考訳(メタデータ) (2023-05-24T15:52:08Z) - Improved generalization by noise enhancement [5.33024001730262]
勾配降下(SGD)の騒音は一般化と密接に関連している。
騒音強調による目標達成手法」を提案する。
その結果,騒音強調による大規模バッチトレーニングは,小バッチトレーニングに比べ,より汎用性が高いことがわかった。
論文 参考訳(メタデータ) (2020-09-28T06:29:23Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。