論文の概要: Visual Attention Faithfulness in Vision-Language Models is Heterogeneous
- arxiv url: http://arxiv.org/abs/2609.00830v1
- Date: Tue, 01 Sep 2026 07:32:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.43247
- Title: Visual Attention Faithfulness in Vision-Language Models is Heterogeneous
- Title(参考訳): 視覚言語モデルにおける視覚的注意の忠実さは不均一である
- Authors: Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo,
- Abstract要約: 視覚的注意の忠実度は不均一であり、3つの異なる処理モードで表されることを示す。
人間の注釈付き接地トラス領域は、モデルアテンションランキングと比較して60ドル%のケースで包括性を満足している。
- 参考スコア(独自算出の注目度): 13.117582974793672
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation analysis on current VLMs, evaluating both the comprehensiveness and sufficiency gap of attention-ranked visual tokens. Our analysis reveals that visual attention faithfulness is heterogeneous, manifesting in three distinct processing modes: Faithful-Sufficient, where top-$k$ attention tokens are both necessary and sufficient for prediction; Faithful-Distributed, where they are necessary but broader visual context remains required; and Non-Focal, where no localized attention region is individually necessary while visual information remains an essential trigger for prediction. Furthermore, human-annotated ground-truth regions satisfy comprehensiveness in only $\sim 60$% of cases compared with model attention rankings, revealing systematic divergence between model visual reliance and human intuition. We demonstrate these patterns across both general VQA on VQAv2 and document tasks on VRDU and ChartQA, showing that visual attention faithfulness varies systematically with processing demands and model architectures rather than being uniformly faithful or unfaithful.
- Abstract(参考訳): 注意重みがモデル推論を忠実に反映するかどうかについては、NLPで活発に議論されているが、視覚言語モデル(VLM)の視覚的モダリティについては、この疑問はほとんど未解決のままである。
本研究は,現在のVLMにおける因果摂動解析を通じて,注目度の高い視覚トークンの包括性と充足的ギャップを両立させることにより,このギャップに対処する。
分析の結果、視覚的注意力の忠実度は3つの異なる処理モードで表されることが明らかとなった。例えば、Fithful-Sufficient, where top-k$ attention tokens is necessary and enough for prediction, Faithful-Distributed, where they is necessary but wide visual context, and Non-Focal, where no localized attention region is required in individually required while visual information is essential trigger for prediction。
さらに、人間の注釈付き接地トラス領域は、モデルアテンションランキングと比較して60ドル程度のケースしかなく、モデルの視覚的依存と人間の直感の体系的な相違が明らかである。
VQAv2上の一般的なVQAとVRDUおよびChartQA上の文書タスクの両方でこれらのパターンを実証し、一様で不誠実ではなく、処理要求やモデルアーキテクチャによって視覚的注意の忠実度が体系的に変化することを示した。
関連論文リスト
- Diagnosing Visual Ignorance in Vision-Language Models [29.2851901986069]
VLM(Vision-Language Models)は、しばしば言語先行に頼り、視覚的証拠に弱く根ざした自信ある答えを生み出す。
本研究では,機械的視点と行動的視点の両方から言語優先性について検討する。
言語優先の信頼性は、モデル内部とベンチマークの妥当性の両方に影響を及ぼす体系的なルーティング障害であることがわかった。
論文 参考訳(メタデータ) (2026-06-05T04:16:01Z) - Unveiling the Visual Counting Bottleneck in Vision-Language Models [49.591496870141846]
この研究は視覚的数え上げを3つの認知段階(視覚的識別、大きさ認識、象徴的マッピング)に分解する。
合成Go基板と線形プローブを用いて、視覚的バックボーンは、外挿系にしっかりと、線形に分離可能な量表現を保っていることを示す。
我々は、崩壊をシンボルマッピングステージに向ける。そこでは、モデルがシンボルトークンに有効な視覚的大きさを投影することに失敗する。
論文 参考訳(メタデータ) (2026-05-28T16:20:29Z) - Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models [84.94288033791346]
我々は,MLLMにおける視覚的表現の劣化という,広範にわたる課題を明らかにするために,詳細な診断分析を行う。
我々は,この現象を,単一のテキスト生成目標によって引き起こされる視覚的犠牲とみなし,そのモデルが解答生成の最適化のためにその視覚的忠実度を損なう。
本研究では,初期視覚特性を予測するために,劣化した中間特徴を強制的に予測し,MLLMの内部表現に固有の視覚特性を維持するための予測正則化を提案する。
論文 参考訳(メタデータ) (2026-03-21T13:10:37Z) - Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models [15.851502442699]
MLLM(Multimodal large language model)はしばしば、拡張推論モードの下で知覚障害に悩まされる。
多段階の推論において、モデルの視覚的注意が散らばり、疑問関連領域から遠ざかって、視覚的入力に効果的に焦点をあてる。
本研究では,エントロピー・フォーカス基準に基づいて視覚的頭部を選択する学習自由な視覚領域誘導注意(VRGA)フレームワークを提案する。
論文 参考訳(メタデータ) (2026-03-15T02:21:05Z) - Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis [21.869968563545736]
我々は、暗黙的な視覚的誤解(IVM)を定義し、MLLMは視覚的入力を完全に理解することなく正しい回答を提供する。
IVMの定量化には,スケール非依存の計量,テクスチャータテンションの精度,新しいベンチマークを導入する。
我々は、より微細な粒度にアプローチを拡張し、その効果を単調なシナリオで実証する。
論文 参考訳(メタデータ) (2025-05-15T17:52:40Z) - VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models [57.43276586087863]
LVLM(Large Vision-Language Models)は幻覚に悩まされ、このモデルでは可聴音を発生させるが、実際には誤出力を発生させる。
既存のベンチマークはスコープに限られており、主にオブジェクト幻覚に焦点を当てている。
対象,属性,関係を多次元のベンチマークで表現し,連想バイアスに基づいて画像を選択する。
論文 参考訳(メタデータ) (2024-04-22T04:49:22Z) - Localization vs. Semantics: Visual Representations in Unimodal and
Multimodal Models [57.08925810659545]
既存の視覚・言語モデルと視覚のみのモデルにおける視覚表現の比較分析を行う。
我々の経験的観察は、視覚・言語モデルがラベル予測タスクに優れていることを示唆している。
我々の研究は、視覚学習における言語の役割に光を当て、様々な事前学習モデルの実証的なガイドとして機能することを願っている。
論文 参考訳(メタデータ) (2022-12-01T05:00:18Z) - Understanding Attention for Vision-and-Language Tasks [4.752823994295959]
本研究では,アテンションスコア計算手法を検討することで,アテンションアライメントの役割を理解するための包括的な分析を行う。
また、注目スコア計算機構がより(あるいはそれ以下)解釈可能な条件も分析する。
我々の分析は,VLタスクの学習段階に適用した場合の,各アテンションアライメントスコア計算の重要性に関する有用な知見を提供する。
論文 参考訳(メタデータ) (2022-08-17T06:45:07Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。