論文の概要: Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
- arxiv url: http://arxiv.org/abs/2606.07604v1
- Date: Fri, 29 May 2026 09:40:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-15 07:09:36.770389
- Title: Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
- Title(参考訳): コントリビューションウェイト:自己注意変換器の幾何学的解析
- Abstract要約: EmphContribution Weightsは、トークンの注意重み、値の大きさ、および層出力との方向アライメントを考慮し、トークンの影響を定量化するプロジェクションベースの計量である。
コントリビューションウェイトはトークンの重要度をより忠実に測定し、意味的に重要なトークンを識別する際の注意ベースの指標を一貫して上回っていることを実証する。
- 参考スコア(独自算出の注目度): 2.211558736382652
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations as it neglects the geometric properties of the value vectors being aggregated. To address this gap, we introduce \emph{Contribution Weights}, a projection-based metric that quantifies a token's influence by accounting for it's attention weight, value magnitude, and directional alignment with the layer output. We demonstrate that contribution weights provide a more faithful measure of token importance, consistently outperforming attention-based metrics in identifying semantically critical tokens across different decoder-only models, tasks, and datasets. Further, our metric enables novel mechanistic analysis of \emph{attention sinks}. While previous work characterized sinks as passive repositories for excess attention, we reveal they serve an active functional role, suppressing information through a convex relationship between sink rate and output norm, stabilizing representations by opposing the semantic drift of low-confidence tokens.
- Abstract(参考訳): 注意重みの分析は,Large Language Models (LLMs) の情報フローを解釈するための標準的なアプローチとなっている。
しかし、このアプローチは、集約される値ベクトルの幾何学的性質を無視するため、かなりの制限がある。
このギャップに対処するために、注視重量、値の大きさ、および層出力との方向アライメントを考慮し、トークンの影響を定量化するプロジェクションベースの計量である \emph{Contribution Weights} を導入する。
コントリビューションウェイトはトークンの重要性をより忠実に測定し、さまざまなデコーダのみのモデル、タスク、データセット間のセマンティッククリティカルトークンの識別において、一貫して注目ベースの指標よりも優れています。
さらに,本測定により,emph{attention sink} の力学解析が可能となった。
過去の研究では、シンクを過剰な注意のためにパッシブレポジトリとして特徴付けていたが、それらはアクティブな機能的役割を担い、シンクレートと出力ノルムの凸関係を通じて情報を抑圧し、低信頼トークンのセマンティックドリフトに反対して表現を安定化させた。
関連論文リスト
- Localized Gaussians as Self-Attention Weights for Point Clouds Correspondence [92.07601770031236]
本稿では,エンコーダのみのトランスフォーマーアーキテクチャのアテンションヘッドにおける意味的意味パターンについて検討する。
注意重みの修正はトレーニングプロセスの促進だけでなく,最適化の安定性の向上にも寄与する。
論文 参考訳(メタデータ) (2024-09-20T07:41:47Z) - Disentanglement via Latent Quantization [60.37109712033694]
本研究では,組織化された潜在空間からの符号化と復号化に向けた帰納的バイアスを構築する。
本稿では,基本データレコーダ (vanilla autoencoder) と潜時再構成 (InfoGAN) 生成モデルの両方に追加することで,このアプローチの広範な適用性を実証する。
論文 参考訳(メタデータ) (2023-05-28T06:30:29Z) - Measuring the Mixing of Contextual Information in the Transformer [0.19116784879310028]
注意ブロック - 複数頭部の注意、残差接続、および層正規化 - を考慮し、トークンとトークンの相互作用を測定するための計量を定義する。
次に,階層的な解釈を集約し,モデル予測のための入力属性スコアを提供する。
実験により,本手法は忠実な説明を提供し,類似のアグリゲーション法より優れていることを示す。
論文 参考訳(メタデータ) (2022-03-08T17:21:27Z) - Attention improves concentration when learning node embeddings [1.2233362977312945]
検索クエリテキストでラベル付けされたノードを考えると、製品を共有する関連クエリへのリンクを予測したい。
様々なディープニューラルネットワークを用いた実験では、注意機構を備えた単純なフィードフォワードネットワークが埋め込み学習に最適であることが示されている。
本稿では,クエリ生成モデルであるAttESTを提案する。このモデルでは,製品とクエリテキストの両方を,潜在空間に埋め込まれたベクトルとして見ることができる。
論文 参考訳(メタデータ) (2020-06-11T21:21:12Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。