論文の概要: Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search
- arxiv url: http://arxiv.org/abs/2609.01431v2
- Date: Thu, 03 Sep 2026 15:16:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-04 16:10:20.257499
- Title: Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search
- Title(参考訳): 電力線エントロピー探索による最適ハイパーパラメータスケーリング法則の効率的な推定
- Authors: Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Bryan Kian Hsiang Low, Eytan Bakshy, Jihao Andreas Lin,
- Abstract要約: マルチフィデリティベイズ最適化に基づくコスト認識獲得関数である Power-Law Entropy Search (PLES) を導入する。
PLESは、単一の目的関数を最適化するのではなく、スケーリング法見積の全体的な不確実性を低減する候補を検索する。
- 参考スコア(独自算出の注目度): 51.33002870475756
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-scale tuning. However, estimating these scaling laws conventionally requires exhaustive grid searches over thousands of training runs, consuming enormous computational resources. We introduce Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates optimal hyperparameter scaling laws through adaptive experimentation. A key innovation in PLES is that it searches for candidates that reduce the overall uncertainty of a scaling law estimate, instead of optimizing a single objective function. At each iteration, PLES selects the candidate configuration that maximally reduces the uncertainty of the scaling law estimates per unit computational cost, naturally favoring informative small-scale experiments. We evaluate PLES on synthetic benchmarks, surrogate models fitted to real LLM training data, and actual LLM pre-training runs. Across all settings, PLES converges to accurate optimal hyperparameter scaling laws using less than one-tenth of the computational budget required by conventional grid search and other baselines.
- Abstract(参考訳): 最適なハイパーパラメータスケーリング法則は、大規模言語モデル(LLM)トレーニングのための最適なハイパーパラメータがモデルとデータスケールでどのように変化するかを記述する。
しかし、これらのスケーリング法則を推定するためには、何千ものトレーニングランに対して網羅的なグリッド探索が必要であり、膨大な計算資源を消費する。
適応実験により最適なハイパーパラメータスケーリング法則を効率よく推定する多要素ベイズ最適化に基づく計算コスト認識獲得関数である Power-Law Entropy Search (PLES) を導入する。
PLESの重要な革新は、単一の目的関数を最適化するのではなく、スケーリング法見積の全体的な不確実性を低減する候補を探すことである。
各イテレーションにおいて、PLESは、単位計算コスト当たりのスケーリング法則推定の不確かさを最大に低減する候補構成を選択する。
実LLMトレーニングデータに適合するサロゲートモデル,および実LLM事前学習実行におけるPLESの評価を行った。
すべての設定において、PLESは従来のグリッドサーチや他のベースラインで必要とされる計算予算の1分の1以下を用いて、正確な最適ハイパーパラメータスケーリング法に収束する。
関連論文リスト
- Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training [7.267441247692648]
本稿では,所定のチェックポイントに対して,計算予算と最適ハイパーパラメータの関係を確立するための新しいフレームワークを提案する。
提案手法は,高パラメータ探索のオーバーヘッドを最大90%削減すると同時に,ベースラインに対して同等あるいは優れた性能を実現する。
このモデルに依存しないフレームワークはアーキテクチャをまたいで一般化し、様々な継続する事前学習シナリオに対して原則的かつ効率的な方法論を提供する。
論文 参考訳(メタデータ) (2026-06-04T02:32:11Z) - Configuration-to-Performance Scaling Law with Neural Ansatz [19.686833161453464]
textitConfiguration-to-Performance Scaling Law (CPL)を学習する
CPLはトレーニング設定が最終トレーニング前損失にどのように影響するかを正確に予測する。
設定に依存しないチンチラ法よりも20~40%低い予測誤差を達成している。
論文 参考訳(メタデータ) (2026-02-10T21:16:59Z) - Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining [59.369484219304866]
我々は100兆のトークンをスクラッチから3,700以上の大規模言語モデル(LLM)に対する前例のない経験的調査訓練を実施している。
ステップ法則(ステップ法)と呼ばれる,LLM事前学習におけるハイパーパラメータ最適化のための普遍的スケーリング法則を確立する。
我々の推定オプティマは, 排他的探索によって得られた世界最高の性能から, テストセットの0.094%しか逸脱しない。
論文 参考訳(メタデータ) (2025-03-06T18:58:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。