論文の概要: Transformers Learn the Mestre-Nagao Heuristic
- arxiv url: http://arxiv.org/abs/2606.15036v1
- Date: Sat, 13 Jun 2026 00:41:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-16 16:21:32.663147
- Title: Transformers Learn the Mestre-Nagao Heuristic
- Title(参考訳): トランスフォーマーはメスト・ナガオ・ヒューリスティックを学ぶ
- Authors: Pranav Venkata Konda,
- Abstract要約: 有理楕円曲線を分類するために、2層変換器エンコーダを訓練する:$E/mathbbQ$ of conductor $leq 10000$ as either rank 0 or rank 1 from the first 128 normalized Frobenius traces。
両クラスで99%の精度を達成し, トレーニングセットにおける等質性や二次的ツイストの相対性を持たない試験曲線では, 基本的に精度は変化しない。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We train a two-layer transformer encoder to classify rational elliptic curves $E/\mathbb{Q}$ of conductor $\leq 10000$ as either rank 0 or rank 1 from the first 128 normalized Frobenius traces. We achieve >99% accuracy on both classes, and accuracy is essentially unchanged on test curves with no isogeny or quadratic-twist relative in the training set. We then apply techniques from mechanistic interpretability such as attention analysis, linear probing, activation patching, logit attribution, and neuron-level circuit analysis to reverse-engineer the algorithm the (centroid in function space) model learned. We find that a sparse circuit of 20 out of 512 layer-1 MLP neurons is sufficient for rank prediction under a linear probe with an AUROC of 0.992 at plateau, implementing a push-pull detector architecture of rank-0 and rank-1 detectors with a one-sided readout. However, we notice that the model has sub-optimal readout problems indicating a mismatch in rank-order between the readout pathway and the discriminative circuit. Critically, the learned input weights of the top discriminating neuron match the Mestre-Nagao sum heuristic weights $\log(p)/(p\cdot \log{B})$ with a Spearman coefficient $r = 0.997$ and Pearson coefficient $r = 0.952$: the model has learnt a result from analytic number theory from the Frobenius trace data alone. We additionally find that all 50 independently trained models concentrate CLS attention on prime positions at 2-50$\times$ the rate of composite positions. The CLS embedding encodes $\log{L(E,1)}$ with $R^2 = 0.962\pm 0.011$ across the 50 models (after controlling for the conductor). Activation patching analysis reveals that attention weights are dissociated from causal information flow. Additionally, the 50 solutions from training are near-identical in function space (with pairwise agreement $>$98.8%) despite large weight space barriers.
- Abstract(参考訳): 有理楕円曲線 $E/\mathbb{Q}$ of conductor $\leq 10000$ as either rank 0 or rank 1 from the first 128 normalized Frobenius traces。
両クラスで99%の精度を達成し, トレーニングセットにおける等質性や二次的ツイストの相対性を持たない試験曲線では, 基本的に精度は変化しない。
次に、注意解析、線形探索、アクティベーションパッチ、ロジット属性、ニューロンレベルの回路解析といった機械論的解釈可能性の手法を適用し、学習した(関数空間におけるセントロイド)モデルを逆エンジニアリングする。
512層1 MLPニューロンのうち20個のスパース回路は、高原で0.992のAUROCを持つ線形プローブの下でのランク予測に十分であり、ランク0およびランク1検出器のプッシュプル検出アーキテクチャを一方の読み出しで実装している。
しかし,本モデルでは,読み出し経路と識別回路のランク順のミスマッチを示す部分最適読み出し問題があることに気付く。
臨界的に、上位判別ニューロンの学習された入力重量は、フロベニウスのトレースデータのみから解析数理論から結果を学んだスピアマン係数 $r = 0.997$ とピアソン係数 $r = 0.952$ の Mestre-Nagao sum Heuristic weights $\log(p)/(p\cdot \log{B})$ と一致する。
さらに、独立に訓練された50のモデルは全て、合成位置のレートが2-50$\times$の素位置に対してCLSの注意を集中していることが分かる。
CLS埋め込みは$\log{L(E,1)}$を$R^2 = 0.962\pm 0.011$でエンコードする(導体を制御した後)。
アクティベーションパッチ解析により、注意重みが因果情報の流れと解離していることが明らかになった。
さらに、トレーニングから得られる50の解は、大きな重み空間障壁にもかかわらず、関数空間においてほぼ同一である(ペアワイズで$98.8%)。
関連論文リスト
- Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index Model [53.6316818897326]
本稿では,2層ニューラルネットワークを時間内にトレーニングするための勾配に基づくアルゴリズムを提案する。
このアルゴリズムは未知の信号$star$と強く一致したスパース表現を学習することを示す。
私たちは、$star$が$k$-sparse for $k = o(sqrtd)$という設定にアプローチを拡張します。
論文 参考訳(メタデータ) (2026-06-13T09:34:39Z) - Efficiently Learning Drifting Halfspaces with Massart Noise [50.4331323695175]
本研究では,マッサートノイズの存在下での漂流概念の学習問題について検討する。
このフレームワークでは、オンライン学習者は独立したサンプルの履歴にアクセスすることができる。
目標は、各ラウンドで小さな予測誤差の仮説を出力することである。
論文 参考訳(メタデータ) (2026-06-09T17:35:18Z) - The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations [50.43168858368539]
大規模言語モデルは自信を持って時代遅れの回答を生成し、既存の方法では検出できない。
これは工学的な失敗ではなく構造的な失敗であり、時間的ドリフトは、幾何的に残留流の方向として、正確性と不確実性の両方に符号化される。
論文 参考訳(メタデータ) (2026-05-09T22:27:31Z) - Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning [61.07540493350384]
自己蒸留(英: Self-distillation, SD)とは、教師自身の予測と地道の混合で学生を訓練する過程である。
任意の予測リスクに対して、各正規化レベルにおいて、最適に混合された学生がリッジ教師に改善されることが示される。
本稿では,グリッド探索やサンプル分割,再構成なしに$star$を推定する一貫したワンショットチューニング手法を提案する。
論文 参考訳(メタデータ) (2026-02-19T17:21:15Z) - Emergence and scaling laws in SGD learning of shallow neural networks [64.48316762675141]
等方性ガウスデータに基づいてP$ニューロンを持つ2層ニューラルネットワークを学習するためのオンライン勾配降下(SGD)の複雑さについて検討した。
平均二乗誤差(MSE)を最小化するために,学生2層ネットワークのトレーニングのためのSGDダイナミックスを高精度に解析する。
論文 参考訳(メタデータ) (2025-04-28T16:58:55Z) - Gradient dynamics for low-rank fine-tuning beyond kernels [9.275532709125242]
学生-教師設定における低ランク微調整について検討する。
基本モデルにおける行列であり,オンライン勾配勾配で訓練された学生モデルが,教師に収束する,という軽微な仮定の下で証明する。
論文 参考訳(メタデータ) (2024-11-23T00:00:28Z) - Generalization and Stability of Interpolating Neural Networks with
Minimal Width [37.908159361149835]
補間系における勾配によって訓練された浅層ニューラルネットワークの一般化と最適化について検討する。
トレーニング損失数は$m=Omega(log4 (n))$ニューロンとニューロンを最小化する。
m=Omega(log4 (n))$のニューロンと$Tapprox n$で、テスト損失のトレーニングを$tildeO (1/)$に制限します。
論文 参考訳(メタデータ) (2023-02-18T05:06:15Z) - Bounding the Width of Neural Networks via Coupled Initialization -- A
Worst Case Analysis [121.9821494461427]
2層ReLUネットワークに必要なニューロン数を著しく削減する方法を示す。
また、事前の作業を改善するための新しい下位境界を証明し、ある仮定の下では、最善を尽くすことができることを証明します。
論文 参考訳(メタデータ) (2022-06-26T06:51:31Z) - An Improved Analysis of Gradient Tracking for Decentralized Machine
Learning [34.144764431505486]
トレーニングデータが$n$エージェントに分散されるネットワーク上での分散機械学習を検討する。
エージェントの共通の目標は、すべての局所損失関数の平均を最小化するモデルを見つけることである。
ノイズのない場合、$p$を$mathcalO(p-1)$から$mathcalO(p-1)$に改善します。
論文 参考訳(メタデータ) (2022-02-08T12:58:14Z) - Online Robust Regression via SGD on the l1 loss [19.087335681007477]
ストリーミング方式でデータにアクセス可能なオンライン環境において、ロバストな線形回帰問題を考察する。
この研究で、$ell_O( 1 / (1 - eta)2 n )$損失の降下は、汚染された測定値に依存しない$tildeO( 1 / (1 - eta)2 n )$レートで真のパラメータベクトルに収束することを示した。
論文 参考訳(メタデータ) (2020-07-01T11:38:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。