論文の概要: Why Adaptive Optimizers Underestimate Rare Tokens
- arxiv url: http://arxiv.org/abs/2609.37535v1
- Date: Tue, 29 Sep 2026 13:20:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.54049
- Title: Why Adaptive Optimizers Underestimate Rare Tokens
- Title(参考訳): アダプティブ・オプティマイザがレアトークンを過小評価する理由
- Abstract要約: アダプティブメソッドは、各更新をその大きさのランニング推定で分割し、トークンが出現した直後にその推定が最大になる。
Adam, Adafactor, Lion, and sign descent はそうではない。
我々は、これらの予測をユニグラムモデルと、既知の生成分布から訓練された小さな言語モデルの両方でテストする。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a running estimate of its magnitude, and that estimate is largest immediately after the token appears. This imbalance has two effects. At the level of the whole output layer, we characterize which optimizers preserve the mean output embedding: every method whose update is linear in past gradients does, as do Kronecker-factored and orthogonalized methods such as Shampoo and Muon. Adam, Adafactor, Lion, and sign descent do not, and for these methods we obtain an exact step-by-step expression for the change. At the level of an individual rare token, the same normalization shifts the training fixed point. In the unigram model, sign descent lowers the logit of every token that occurs in fewer than half of the minibatches at a constant expected rate. For RMSProp with periodic arrivals, we can solve the fixed point in closed form: if a token is absent for at least two consecutive minibatches, its equilibrium probability is strictly below its data frequency for every learning rate, and the ratio tends to $κ/(2(e^{κ/2}-1))$. Here $κ$ is the mean number of steps between occurrences divided by the second-moment time constant $1/(1-β_2)$. In the same model, SGD and AMSGrad retain the unbiased fixed point. We test these predictions both in a unigram model and in a small language model trained from a known generating distribution. With random arrivals, the bias is larger than the periodic formula predicts; in the language model, the optimizers with the biased fixed point also fit the generating distribution less well.
- Abstract(参考訳): ソフトマックス出力層では、希少なトークンは、ほとんどのステップで小さな正のロジット勾配を受け取り、ターゲットである時には少数のステップではるかに大きな負の勾配を受ける。
SGDはこれらの貢献を単純に追加する。
Adam, RMSProp, sign descent などのコーディネート・ワイド・アダプティブな手法は、代わりに各更新をその大きさのランニング推定で割る。
この不均衡には2つの効果がある。
出力層全体のレベルでは、どのオプティマイザが平均出力埋め込みを保存するかを特徴付けます。
Adam, Adafactor, Lion, and sign descent はそうではない。
個々のレアトークンのレベルでは、同じ正規化がトレーニング定点をシフトする。
一グラムモデルでは、符号降下はミニバッチの半分以下で起こる全てのトークンのロジットを一定の期待レートで低下させる。
周期的到着を持つ RMSProp に対して、固定点を閉じた形で解くことができる: トークンが少なくとも2つの連続するミニバッチに対して欠如している場合、その平衡確率は学習速度毎にそのデータ周波数より厳密に低く、その比率は$κ/(2(e^{κ/2}-1))$である。
ここで、κ$は第2モーメント時間定数1/(1-β_2)$で割った事象間の平均ステップ数である。
同じモデルでは、SGD と AMSGrad は偏りのない固定点を保持する。
我々は、これらの予測をユニグラムモデルと、既知の生成分布から訓練された小さな言語モデルの両方でテストする。
ランダム到着の場合、バイアスは周期式よりも大きく、言語モデルでは、バイアスのある固定点を持つオプティマイザも生成する分布に適合しない。
関連論文リスト
- On the Two Faces of Adam in Separable Linear Classification [30.031267910132218]
ログロス下でのソフトマックスパラメトリゼーションを用いた分離線形分類において, 決定論的, 完全バッチ, バイアス補正されたAdamの挙動を考察する。
アダムは安定性定数$$が 0 であるとき、最大ノルム・マルジン最適度に近付くことが知られ、一方正の$はユークリッド・マルジン最適度に近付くことが知られている。
論文 参考訳(メタデータ) (2026-09-27T20:37:36Z) - Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences [17.927451930317197]
自己回帰言語モデルの実証的なスケーリング法則は、予測損失をモデルサイズ、データサイズ、最適化計算に関連付ける。
本研究では,安定な潜在線形RNNが軌道を発生し,スケッチされた線形リカレント学習者は,次の予測に基づいて,安全でフルバッチなWSD勾配勾配で学習する,学習可能な教師-学生モデルを用いて検討する。
論文 参考訳(メタデータ) (2026-08-23T13:37:15Z) - What Does a Discrete Diffusion Model Learn? [71.03603607324338]
離散拡散モデルは、デノイザ、スコア比、ブリッジプラグイン予測器などを学ぶ。
まず, 連続時間マルコフ連鎖 (CTMC) ELBO の任意のノイズ発生過程に対する厳密な導出から始める。
すべてのアイデンティティは、正確に解けるモデル上で近似なしで数値的に検証される。
論文 参考訳(メタデータ) (2026-07-06T17:56:11Z) - Information Hidden in Gradients of Regression with Target Noise [2.8911861322232686]
勾配だけでヘッセンが明らかになることを示す。
我々はガウス以下の入力の下で非漸近作用素ノルム保証を提供する。
論文 参考訳(メタデータ) (2026-01-26T14:50:16Z) - Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification [50.717692060500696]
対数損失を伴う次のトーケン予測は自己回帰シーケンスモデリングの基盤となる。
次トーケン予測は、適度な誤差増幅を表す$C=tilde O(H)$を達成するために堅牢にすることができる。
C=e(log H)1-Omega(1)$。
論文 参考訳(メタデータ) (2025-02-18T02:52:00Z) - Second-order Information Promotes Mini-Batch Robustness in Variance-Reduced Gradients [0.196629787330046]
目的関数の部分的な2次情報を組み込むことで、分散還元勾配法のミニバッチサイズに対するロバスト性を劇的に向上させることができることを示す。
本稿では,この現象をプロトタイプNewton(textttMb-SVRN$)アルゴリズムで示す。
論文 参考訳(メタデータ) (2024-04-23T05:45:52Z) - Precise Learning Curves and Higher-Order Scaling Limits for Dot Product
Kernel Regression [41.48538038768993]
本稿では,ドット積カーネルのカーネルリッジ回帰問題に焦点をあてる。
我々は、任意の整数$r$に対して$m approx dr/r!$が常に学習曲線のピークを観測し、複数のサンプルワイズと非自明な振る舞いを複数のスケールで達成する。
論文 参考訳(メタデータ) (2022-05-30T04:21:31Z) - High-dimensional Asymptotics of Feature Learning: How One Gradient Step
Improves the Representation [89.21686761957383]
2層ネットワークにおける第1層パラメータ $boldsymbolW$ の勾配降下ステップについて検討した。
我々の結果は、一つのステップでもランダムな特徴に対してかなりの優位性が得られることを示した。
論文 参考訳(メタデータ) (2022-05-03T12:09:59Z) - Correcting Momentum with Second-order Information [50.992629498861724]
最適積に$O(epsilon)$epsilon点を求める非臨界最適化のための新しいアルゴリズムを開発した。
我々は、さまざまな大規模ディープラーニングベンチマークとアーキテクチャで結果を検証する。
論文 参考訳(メタデータ) (2021-03-04T19:01:20Z) - Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and
Variance Reduction [63.41789556777387]
非同期Q-ラーニングはマルコフ決定過程(MDP)の最適行動値関数(またはQ-関数)を学習することを目的としている。
Q-関数の入出力$varepsilon$-正確な推定に必要なサンプルの数は、少なくとも$frac1mu_min (1-gamma)5varepsilon2+ fract_mixmu_min (1-gamma)$の順である。
論文 参考訳(メタデータ) (2020-06-04T17:51:00Z) - Carath\'eodory Sampling for Stochastic Gradient Descent [79.55586575988292]
本稿では,Tchakaloff と Carath'eodory の古典的な結果から着想を得た手法を提案する。
我々は、測定値の低減を行う降下ステップを適応的に選択する。
これをBlock Coordinate Descentと組み合わせることで、測定の削減を極めて安価に行えるようにします。
論文 参考訳(メタデータ) (2020-06-02T17:52:59Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。