論文の概要: Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
- arxiv url: http://arxiv.org/abs/2605.04893v1
- Date: Wed, 06 May 2026 13:25:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-07 18:41:07.834383
- Title: Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
- Title(参考訳): 移動としての自己注意:対称スペクトル診断の限界
- Authors: Dominik Dahlem, Diego Maniloff, Mac Misiura,
- Abstract要約: 均一な因果的注意は、$n$に依存しないフロア$ge 1/5$で、最悪のカットは$tast/n approx 0.32$で、ウィンドウアテンションは$O(w/n)$で、フロアは$O(w/n)$である。
結果として生じる2軸の診断は、ファルシブルな極性予測をもたらす。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models hallucinate in predictable ways: attention routing fails by over-concentrating on a narrow set of positions, or by spreading so diffusely that relevance is diluted, and the shape of the failure carries diagnostic signal. A widely used family of spectral methods analyzes the symmetric component of the degree-normalized attention operator, which governs transport capacity; we prove that every transpose-invariant spectral diagnostic of this operator is structurally orientation-blind (it cannot distinguish an operator from its transpose, and therefore cannot detect information-flow direction), with a quantitative converse establishing the asymmetry coefficient $G$ as the unique control parameter for direction. Pairing this with a closed-form bipartite-Cheeger landscape for canonical causal architectures, we show that uniform causal attention satisfies an $n$-independent floor $φ\ge 1/5$ with worst cut at $t^\ast/n \approx 0.32$, while window attention pierces the floor as $O(w/n)$; failure modes are shape-different, not just value-different. The resulting two-axis diagnostic ($φ$ for capacity, $G$ for direction) yields a falsifiable polarity prediction: bottleneck- and diffuse-dominated benchmarks should exhibit opposite polarity. Under length-controlled evaluation, transport features retain interpretable signal (LC-AUROC from 0.62 to 0.84) on tested models up to 8B parameters, with polarity reversing as predicted between HaluEval and MedHallu.
- Abstract(参考訳): 大きな言語モデルは、予測可能な方法で幻覚を与える: 注意ルーティングは、狭い位置の集合に過度に集中することによって失敗するか、あるいは、関係が希薄になり、障害の形状が診断信号を運ぶように拡散することによって失敗する。
スペクトル法の一群は,輸送能力を管理する次数正規化アテンション演算子の対称成分を解析し,この演算子のすべての転位不変変分診断が構造的配向盲であること(演算子と転位を区別できず,従って情報フロー方向を検出することができない)を,非対称係数を方向のユニークな制御パラメータとして定量的に定式化して解析する。
これを正準因果アーキテクチャのための閉形式バイパートイト・シーガーランドスケープと組み合わせることで、均一な因果注意が$n$独立フロア$φ\ge 1/5$で最悪のカットが$t^\ast/n \approx 0.32$であるのに対して、ウィンドウアテンションは$O(w/n)$である。
結果として生じる2軸診断(キャパシティはφ$、方向はG$)は、偽の極性予測をもたらす。
長さ制御による評価では、輸送特性は8Bパラメータの試験モデル上で解釈可能な信号(LC-AUROC:0.62から0.84まで)を保持し、HaluEvalとMedHalluの予測通り極性反転を行う。
関連論文リスト
- Cross-Spectral Witness for Hidden Nonequilibrium Beyond the Scalar Ceiling [18.923595971721344]
粗粒化は、明らかに平衡のような還元された記述に隠された強制を吸収する可能性がある。
線形系のスカラーオブザーバブルでは、時間可逆性統計学が基礎となる駆動を検出できない。
一般的なCSMでは、共有シークタードライブを認証する。
論文 参考訳(メタデータ) (2026-04-04T15:54:29Z) - Stability and Generalization of Push-Sum Based Decentralized Optimization over Directed Graphs [55.77845440440496]
プッシュベースの分散通信は、情報交換が非対称である可能性のある通信ネットワークの最適化を可能にする。
我々は、グラディエント・プッシュ(SGP)アルゴリズムのための統一的な一様安定性フレームワークを開発する。
重要な技術的要素は、2つの量に束縛された不均衡認識の一般化である。
論文 参考訳(メタデータ) (2026-02-24T05:32:03Z) - Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations [57.179679246370114]
乱摂動の分布は, 摂動段差がゼロになる傾向にあるため, 推定子の分散を最小限に抑える。
以上の結果から, 一定の長さを維持するのではなく, 真の勾配に方向を合わせることが可能であることが示唆された。
論文 参考訳(メタデータ) (2025-10-22T19:06:39Z) - Transformers as Support Vector Machines [54.642793677472724]
自己アテンションの最適化幾何と厳密なSVM問題との間には,形式的等価性を確立する。
勾配降下に最適化された1層変圧器の暗黙バイアスを特徴付ける。
これらの発見は、最適なトークンを分離し選択するSVMの階層としてのトランスフォーマーの解釈を刺激していると信じている。
論文 参考訳(メタデータ) (2023-08-31T17:57:50Z) - Neural-network quantum state study of the long-range antiferromagnetic Ising chain [0.771303749110121]
横磁場イジング鎖の反強磁性相互作用を代数的に減衰させた反強磁性相互作用における量子相転移について検討する。
SR極限の普遍比が$alpha_mathrmLR 2$で成り立たないことが、臨界度の偏りを示唆している。
論文 参考訳(メタデータ) (2023-08-18T17:58:36Z) - Analytic Signal Phase in $N-D$ by Linear Symmetry Tensor--fingerprint
modeling [69.35569554213679]
解析信号位相とその勾配は2-D$以上の不連続性を持つことを示す。
この欠点は深刻なアーティファクトをもたらす可能性があるが、問題は1-D $シグナルには存在しない。
本稿では,複数のGaborフィルタに頼って線形シンメトリー位相を用いることを提案する。
論文 参考訳(メタデータ) (2020-05-16T21:17:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。