論文の概要: What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
- arxiv url: http://arxiv.org/abs/2607.12735v1
- Date: Tue, 14 Jul 2026 13:06:44 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-15 17:08:30.151433
- Title: What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
- Title(参考訳): 表現に先行する作業とは何か? 特徴ファミリ, ラベルのない不変性, グローキングにおける重要なWindows
- Authors: Gunner Levi Howe,
- Abstract要約: 我々は、4つの軸にまたがって、188の新たなランで、そのような先行的な作業を特徴付ける。
間違ったフィーチャーファミリから構築された一貫性のある学習可能な事前ビルドは、ランダムパーティションのように一般化をブロックする。
前者は早期に必要であり、最初の2000年代のみにのみ適用される。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learnable prior built from the wrong feature family (magnitude bands) blocks generalization like a random partition (1/15 vs 0/20 grok; $p=0.43$ between them), confirming the companion's prediction that priors act at the level of the circuit's features. Supervision: a fully label-free invariance prior -- positives are commuted pairs $(a,b)\sim(b,a)$ only -- generalizes in 15/15 runs at a median $2.7\times$ speedup, more reliably than the label-supervised prior itself ($p=0.038$), and combined with a weight-norm clamp yields the strongest method we test (median $17\times$, 5/5) -- strongest meaning reliably fast: plain cross-entropy with a clamp matches this speed only at the exact critical norm, while the prior keeps it fast across the entire clamp range. Timing: the prior is only needed early -- applied solely during the first 2000 epochs (4% of budget) it generalizes 10/10 at $2.7\times$, beating continuous application (8/10, $1.25\times$) and a duration-matched later window ($2.1\times$). Setting: the dissociation replicates on modular multiplication and across depths and normalization variants, and a clamp sweep quantifies the companion's central claim: structure injection flattens the weight-norm delay-law exponent about 17-fold (plain cross-entropy slows $31\times$ per +10 norm units, a lower bound as higher cells are censored, versus $1.22\times$ with the prior). Honest boundary: tasks that generalize before memorizing have no delay to control. Feature-family alignment decides whether a prior permits generalization; invariance content suffices for acceleration without labels; a brief early window captures nearly all of the benefit.
- Abstract(参考訳): コンパニオン・ワークは、グルーキング遅延が、対照的に先行して注入可能なタスク構造化表現を形成するための時間であることを示した。
ここでは、4つの軸にまたがって、188の新たなランで、そのような先行的な作業を特徴付ける。
コンテンツ: 間違ったフィーチャーファミリ(マグニチュードバンド)から構築されたコヒーレントで学習可能な事前構築は、ランダムなパーティション(1/15 vs 0/20 grok; $p=0.43$)のような一般化をブロックし、先行がサーキットの特徴のレベルに作用するという仲間の予測を確認する。
Supervision: complete label-free invariance prior -- positives are commuted pairs $(a,b)\sim(b,a)$ only -- generalizes in 15/15 run at a Medium $2.7\times$ speedup, more reliable to the label-supervised prior itself (p=0.038$) and combined with a weight-norm clamp yields the most method (median $117\times$, 5/5) -- strong meaningably fast: plain cross-entropy with clamp with a clamp match this speed at the exact critical norm, while the prior keep it keep fast across the clamp range.
タイミング: 前者は早期に必要であり、最初の2000年代(予算の4%)にのみ適用される。これは10/10を2.7\times$で一般化し、連続的なアプリケーション(8/10, $1.25\times$)を圧倒し、後続のウィンドウ(2.1\times$)を延長する。
構造注入は17倍の重量-ノルム遅延則指数を平坦化させる(平らなクロスエントロピーは311\times$ +10ノルム単位あたり311\times$を遅くし、より高いセルが検閲されるにつれて、より低いバウンドは1.22\times$である)。
正直な境界: 記憶する前に一般化するタスクは制御に遅れない。
特徴系列アライメントは、前者が一般化を許すかどうかを決定し、不変コンテンツはラベルなしで加速するのに十分である。
関連論文リスト
- Structure-Specific Representational Priors Causally Control the Grokking Delay [0.0]
グローキングは、トレーニングセットからずっと後の一般化であり、構造に依存しない介入によって加速されている。
ラベルではなく機能レベルで決定された、適切な表現構造を形成するのは慎重なタイミングであることを示す。
加速度はウェイトノームのサイドエフェクトによって促進されるので、トレーニング中にノルムをクランプすると、信頼できるスタンドアロンの加速器が得られる。
論文 参考訳(メタデータ) (2026-07-05T14:32:26Z) - Attention is Just Another Name for Coupling?: A Fast-Slow ODE Perspective on Hierarchical Pretraining [0.0]
本稿では,ゼロ初期化ゲート補間により,時間的に遅い第2のカップリングが高速経路にフィードバックされるかどうかを問う。
本論文は、高速スローODE形式を具体的なニューラルネットワークとしてインスタンス化する。
論文 参考訳(メタデータ) (2026-06-15T13:54:15Z) - Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge [0.0]
本稿では,TPUハードウェア上での長期的コンテキストバランスによる最適輸送注意度を,停止ベース,固定深部テールリファインメントサロゲートを用いて検討する。
本稿では, 局所的代理バイアスバウンド, 後部バイアス証明書, および, 厳密な正の能動ブロックに対する射影収縮証明書を提供する。
合成マスク問題では、最適化された置換基の正確なオートディフは10〜5ドル-10〜10ドルである。
論文 参考訳(メタデータ) (2026-04-28T19:01:11Z) - Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models [85.18197548789291]
PRISM-$ (Projection-based Relevance-Informed Steering Method) を提案する。
PRISM-$$は20種中19種で最も優れた方法であり、相対的な利得は+10.6%まで上昇し、ステアリングのコストは半減する。
PRISM-$はFlashAttentionと互換性があり、無視できるメモリオーバーヘッドを追加する。
論文 参考訳(メタデータ) (2026-03-11T12:24:45Z) - Robust Layerwise Scaling Rules by Proper Weight Decay Tuning [50.11170157029911]
現代のスケール不変アーキテクチャでは、トレーニングは急速に劣化したグラデーション状態に入る。
我々は,AdamWに対して,幅をまたいだサブ層ゲインを保ったウェイトデカイスケーリングルールを導入する。
この結果は,パラメータが設定した定常スケールを明示的に制御することにより,ほぼ入出力体制を超えて$mu$Pを拡大する。
論文 参考訳(メタデータ) (2025-10-17T02:58:35Z) - Proving the Limited Scalability of Centralized Distributed Optimization via a New Lower Bound Construction [57.93371273485736]
我々は、すべての労働者が同一の分布にアクセスする均質な(すなわちd.d.)場合であっても、すべての労働者が非バイアス付き境界 LDeltaepsilon2,$$$$$ のポリ対数的により良いポリ対数を求める集中型分散学習環境を考える。
論文 参考訳(メタデータ) (2025-06-30T13:27:39Z) - From Continual Learning to SGD and Back: Better Rates for Continual Linear Models [50.11453013647086]
以前見られたタスクの損失を、$k$の繰り返しの後、忘れること、すなわち、分析する。
実現可能な最小二乗の設定において、新しい最上界を創出する。
我々は、タスクを繰り返しないランダム化だけで、十分に長いタスクシーケンスで破滅的な事態を防げることを初めて証明した。
論文 参考訳(メタデータ) (2025-04-06T18:39:45Z) - Sharper Convergence Guarantees for Asynchronous SGD for Distributed and
Federated Learning [77.22019100456595]
通信周波数の異なる分散計算作業者のトレーニングアルゴリズムを示す。
本研究では,より厳密な収束率を$mathcalO!!(sigma2-2_avg!)とする。
また,不均一性の項は,作業者の平均遅延によっても影響されることを示した。
論文 参考訳(メタデータ) (2022-06-16T17:10:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。