論文の概要: Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?
- arxiv url: http://arxiv.org/abs/2610.04119v1
- Date: Fri, 02 Oct 2026 22:50:59 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 21:04:30.017511
- Title: Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?
- Title(参考訳): 抑うつ前のコピー:言語モデルトレーニング中に低速度のディップを駆動するものは何か?
- Abstract要約: ピシアのモデルは初期の訓練窓を通過し、繰り返し名前が好まれる。
繰り返し名前を書くアテンションヘッドのセットを選択することで、誤った選好の原因を1つ特定する。
これにより、個別に訓練された160Mモデルと公式の160M、410M、1Bモデルにおいて、正しい最小繰り返しのロジット差が改善される。
- 参考スコア(独自算出の注目度): 0.0
- License:
- Abstract: Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass through an early training window in which they prefer the repeated name, so accuracy in a choice between the two names falls below one half while language-model loss on a fixed text sample keeps decreasing across the same window. The window reflects a temporary imbalance between two computations. We identify one cause of the wrong preference by selecting a set of attention heads that write the repeated name, on prompts separate from those used for causal evaluation, keeping that selection fixed, and then replacing each head's final-token output with its average output on a separate set of non-repeated-name prompts. This improves the correct-minus-repeated logit difference in a separately trained 160M model and in the official 160M, 410M, and 1B models. At 160M, the head that lowers the repeated name in the mature model shows little of its mature behaviour at this point. It directs less than one percent of its attention to the repeated mention, and its output makes almost no direct contribution to lowering that name's logit. Both properties grow over the interval in which behaviour recovers. Across the 160M, 410M, and 1B models, transplanting the corresponding head's mature parameters into the early checkpoint recovers 35 to 68 percent of the total improvement in the correct-minus-repeated logit difference seen by the end of training. Related early-to-late reversals appear at further Pythia scales, in two independently trained GPT-2 models, and in OLMo. A mature circuit can therefore conceal a transient causal configuration that shaped behaviour earlier in training.
- Abstract(参考訳): 機械的解釈可能性は通常、完全に訓練されたモデルを研究するが、モデルがまだタスクを学習している間に、振る舞いを駆動する計算は変化する可能性がある。
間接オブジェクト識別タスクでは、モデルが2度言及された名前ではなく、1度言及された名前を使い続けるべきです。
Pythiaモデルは、繰り返し名前を好む初期のトレーニングウィンドウを通過するので、2つの名前の選択の精度は半分以下になり、固定されたテキストサンプルの言語モデル損失は、同じウィンドウをまたいで減少し続けている。
ウィンドウは2つの計算間の一時的な不均衡を反映する。
我々は、繰り返し名前を書くための注意ヘッドのセットを選択し、因果評価に用いるものと異なるプロンプトで選択し、その選択を固定し、その後、各ヘッドの最終的な出力を、別の非繰り返し名前プロンプトのセットで平均出力に置き換えることで、誤った好みの1つの原因を特定する。
これにより、個別に訓練された160Mモデルと公式の160M、410M、1Bモデルにおいて、正しい最小繰り返しのロジット差が改善される。
160Mでは、成熟したモデルで繰り返し名前を下げる頭部は、現時点ではその成熟した振る舞いをほとんど示していない。
繰り返し言及に注意を向ける割合は1%に満たず、その出力はその名前のロジットを下げるための直接的な貢献はほとんどない。
両方の特性は、振る舞いが回復する間隔で成長する。
160M、410M、および1Bモデル全体で、対応する頭部の成熟したパラメータを早期チェックポイントに移植すると、トレーニングの終了によって見られる正しい最小繰り返しロジット差の合計改善の35~68%が回復する。
関連する早期から後期の逆転は、Pythiaスケール、独立に訓練された2つのGPT-2モデル、OLMoに現れる。
したがって、成熟した回路は、トレーニングの早い段階での振る舞いを形作る過渡的な因果構成を隠蔽することができる。
関連論文リスト
- Double Descent and Malign Overfitting in Diffusion Models [5.939780039158003]
拡散モデルの過度な適合は破滅的であり、モデルを体制へと駆り立てる。
トレーニングサンプルあたりの雑音実現の固定数$m$では、2次ピークが発生するが、標準回帰のように$psim n$ではなく$psim nm$となる。
この過度な適合は、トレーニングの暗黙の規則化が完全に作業中であるにもかかわらず、真のスコアではなく、経験的なスコアに向かってモデルを駆動するからである。
論文 参考訳(メタデータ) (2026-09-22T13:30:48Z) - Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans [1.2109382221827507]
5つのモデルファミリーにわたる15のモデルに反復プライミングを適用する。
ベースモデルが自動処理を示し,インストラクションモデルが制御処理を示すことがわかった。
人間はハイブリッドなプロファイルを示し、ラグ感受性のファシリテーションはインストラクションモデルに似ているが干渉はしない。
論文 参考訳(メタデータ) (2026-08-05T00:52:03Z) - Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training [57.771940716189114]
我々は、大きな言語モデル(LLM)が「逆の呪い」に苦しむことを示す。
逆の呪いの根本原因は、訓練と推論の段階で異なる単語順にある。
この問題に対処するために,SPT(Semantic-Aware Permutation Training)を提案する。
論文 参考訳(メタデータ) (2024-03-01T18:55:20Z) - Training Trajectories of Language Models Across Scales [99.38721327771208]
言語モデルのスケールアップは、前例のないパフォーマンス向上につながった。
異なるサイズの言語モデルは事前学習中にどのように学習するか?
より大きな言語モデルはなぜ望ましい振る舞いを示すのか?
論文 参考訳(メタデータ) (2022-12-19T19:16:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。