論文の概要: Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- arxiv url: http://arxiv.org/abs/2607.20524v1
- Date: Thu, 09 Jul 2026 07:04:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-27 00:46:13.217117
- Title: Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Title(参考訳): 大規模言語モデルにおける注意欠陥,機能トークンアンチョリング,および注意に基づく介入の限界
- Authors: Sagar Dangal, Manoj Shakya,
- Abstract要約: GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, distilgpt2の6つの協調実験を行った。
まず,短期(5-100トークン)の注意低下を特徴付ける。
注意力の低下は規範的というよりはむしろ典型的であると結論付けている。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested. We present six coordinated experiments across GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, and distilgpt2. We first characterise short-term (5-100 token) attention degradation, finding a universal exponential-then-plateau pattern whose rate is inversely correlated with depth, with distinct layer-wise entropy signatures per architecture. Function token anchoring proves architecture-dependent: OPT-1.3B (absolute positional encoding) shows distance-dependent preposition specificity, GPT-2 shows uniform non-specific dependence, and LLaMA (RoPE) shows reversal at long distances. Strategic comma insertion at clause boundaries causally reduces prediction degradation in the 40-80 token range, with the benefit tied to syntactic boundary alignment rather than token density. We then test the mechanism causally: Relay-Aware Attention (RAA), which biases attention logits toward function token positions, verifiably increases attention mass by 16-24% yet yields null effects on GPT-2 and LLaMA-1B, preliminary harm on LLaMA-3B, and a mixed effect on OPT-1.3B that nets to approximately zero. Multi-fact retrieval probes further show that degradation rate does not predict retrieval accuracy across models. We conclude that mean attention degradation is largely descriptive rather than prescriptive: function tokens contribute through what their hidden states compute, not through the attention they receive -- with implications for interpretability methodology and attention-score-based inference optimisations such as KV-cache eviction.
- Abstract(参考訳): 意味的横断的アテンション劣化はトランスフォーマーの解釈可能性において広く報告されているが、文脈的検索を因果的に制限するか否かは未検証のままである。
GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, distilgpt2の6つの協調実験を行った。
まず,短時間(5-100トークン)のアテンション劣化を特徴付け,各構造ごとに異なる層次エントロピーシグネチャで,その速度が深さと逆相関する万能指数テンプラトーパターンを求める。
OPT-1.3B(絶対位置エンコーディング)は距離依存のプリポジション特異性を示し、GPT-2は均一な非特異性を示し、LLaMA(RoPE)は長距離で反転を示す。
節境界における戦略的コマ挿入は、トークン密度よりも構文的境界アライメントに結びついているため、40-80トークン範囲の予測劣化を因果的に減少させる。
注意ログを関数トークンの位置にバイアスするRAA(Relay-Aware Attention)は、注意質量を16~24%増加させるが、GPT-2とLLaMA-1Bには無効効果、LLaMA-3Bには予備的障害、OPT-1.3Bにはほぼゼロに混合効果を与える。
マルチファクト検索プローブは、劣化率がモデル間での精度を予測しないことを示す。
関数トークンは、認識可能性の方法論やKV-cache消去のような注目スコアに基づく推論の最適化に影響を及ぼす。
関連論文リスト
- SALT-GNN: Handling Dense Neighborhoods in Anti-Money Laundering Graphs via Statistics-Aware Attention [7.487248869469991]
マネーロンダリングは金融の安定を脅かし、機関に罰則を課し、自動検出を動機付けている。
レイダリングスキームはリレーショナルパターンを通じてしばしば現れるため、グラフニューラルネットワーク(GNN)は反マネーロンダリング(AML)にますます利用されている。
本稿では,各メッセージパッシング層に注意を払って,次数対応の統計アグリゲーションを融合する統計アグリゲーションアーキテクチャSALT-GNNを提案する。
論文 参考訳(メタデータ) (2026-07-11T05:42:19Z) - When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection [12.446807294893638]
本稿では,視線方向の相互コヒーレンス,頭部アライメント,対人関係の瞳孔配置として定義された高レベルの意味的キューであるソーシャル・ゲイズ・コンシステンシーを紹介する。
既存の低レベルパラダイムに対して,これまで未利用であった検出軸を構成することを示す。
4ステップのアカウントでは、単一インパインター(FLUX.1-Fill)のトレーニングがマルチジェネレータスイートに移行した理由が説明されている。
論文 参考訳(メタデータ) (2026-05-26T17:50:17Z) - Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits [1.840562129212051]
広範にわたる見解では、可聴言語モデル(VLM)は、注意が鋭いときに最も信頼できるものである。
注意構造、生成ダイナミクス、隠れ状態幾何を1つの正しさラベルと比較する。
3-7BのVLMでは、アテンションマップのシャープネスよりも、隠れ状態の幾何、層幅のマージン形成、スパースレイト層の回路から信頼性を確実に読み取ることができる。
論文 参考訳(メタデータ) (2026-05-05T22:27:05Z) - Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation [55.74376789006731]
視覚的インストラクションデータに対するSFT(Supervised Fine-tuning)は、しばしば視覚言語モデル(VLM)の知覚能力を向上し、推論性能を低下させる。
この劣化が深度表現の障害的アクセスと関係しているかどうかを考察し、固定された深度集合でさえ推論を著しく復元することを示した。
IADAは、クロスディープな入力適応性、モダリティを意識し、低ランクのボトルネックを通じて効率的にパラメータ化できる軽量なメカニズムである。
論文 参考訳(メタデータ) (2026-03-27T11:47:39Z) - Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models [44.28116882776357]
textbfPunctuation-aware textbfHybrid textbfSparse textbfAttention textbf(PHSA)を提案する。
具体的には,大域的セマンティック表現と句読点付き境界特徴を融合させ,コアセマンティック構造を保ちながら,計算オーバーヘッドをほとんど含まない二重ブランチアグリゲーション機構を設計する。
論文 参考訳(メタデータ) (2026-01-06T08:47:16Z) - DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models [55.30555646945055]
テキスト・ツー・イメージ(T2I)モデルはセマンティック・リークに対して脆弱である。
DeLeakerは、モデルのアテンションマップに直接介入することで、漏洩を緩和する軽量なアプローチである。
SLIMはセマンティックリークに特化した最初のデータセットである。
論文 参考訳(メタデータ) (2025-10-16T17:39:21Z) - vAttention: Verified Sparse Attention [100.98210818821688]
vAttentionは、ユーザが指定した$(epsilon, delta)$の近似精度保証(thus, confirmed)を備えた実用的なスパースアテンションメカニズムである。
vAttentionはデータセット間のスパースアテンションの質を大幅に改善することを示す。
モデルの品質を損なうことなく高速なデコードを実現するために、推論シナリオにデプロイすることができる。
論文 参考訳(メタデータ) (2025-10-07T08:46:08Z) - Ladder-of-Thought: Using Knowledge as Steps to Elevate Stance Detection [73.31406286956535]
姿勢検出タスクにLadder-of-Thought(LoT)を導入する。
LoTは、小さなLMに高品質な外部知識を同化させ、生成した中間的論理を精査するように指示する。
実験では, 姿勢検出タスクにおけるCoTのGPT-3.5よりも16%改善し, 10%向上した。
論文 参考訳(メタデータ) (2023-08-31T14:31:48Z) - STAR Loss: Reducing Semantic Ambiguity in Facial Landmark Detection [80.04000067312428]
本稿では,意味的あいまいさの特性を利用した自己適応型あいまいさ低減(STAR)の損失を提案する。
意味的あいまいさは異方性予測分布をもたらすことが分かり、予測分布を用いて意味的あいまいさを表現する。
また,分布の異常変化とモデルの初期収束を回避できる2種類の固有値制限法を提案する。
論文 参考訳(メタデータ) (2023-06-05T10:33:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。