論文の概要: Disentangling Semantic Attention from Structural Bias in the Attention Manifold
- arxiv url: http://arxiv.org/abs/2607.24017v1
- Date: Mon, 27 Jul 2026 05:30:19 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.322297
- Title: Disentangling Semantic Attention from Structural Bias in the Attention Manifold
- Title(参考訳): アテンション・マニフォールドにおける構造バイアスからのセマンティック・アテンションの遠ざかる
- Abstract要約: MLLM(Multimodal Large Language Models)は、意味的に非形式的な視覚トークンに対して不均等な注意を払っている。
本研究では,SPAR(Saliency-guided Purification and Adaptive Redistribution)を導入した。
- 参考スコア(独自算出の注目度): 61.68044165974505
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.
- Abstract(参考訳): MLLM(Multimodal Large Language Models)における注意機構の実証的な成功は、その固有の微妙な欠陥を曖昧にすることが多い。
具体的には、MLLMは、特定の意味的に非形式的な視覚トークンに対して、常に不均等な注意を払っている。
既存の推論干渉法は、これらのシンクトークンを識別し、注意重みを再分配しようとするが、そのような手法は通常、これらのトークンを分離して扱い、計算の非効率さに悩まされる。
代わりに、我々はこの現象を、孤立したシンクトークンを超えて広がる視覚的特徴に課せられる一般化されたテキストバイアスとして再考した。
この観点から、広範にわたる構造バイアスは、意味的視覚信号の希釈につながり、モデルが有効な視覚的証拠よりも言語的先行を優先するにつれて、多モーダル幻覚を誘発する。
この制限に対処するために、SPAR(Saliency-guided Purification and Adaptive Redistribution)を導入し、トレーニング不要でプラグアンドプレイの介入を行う。
SPARは、この一般化されたテキストバイアスを、構造的ノイズを浄化し、再生された注意予算を最も情報性の高い視覚領域に再分配することで緩和する。
様々な幻覚ベンチマークの包括的評価は、SPARが無視可能な計算オーバーヘッドで真の視覚的グラウンドを効果的に復元することを証明している。
関連論文リスト
- ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs [77.57757084153083]
本稿では,効率的なMLLMを実現するために,Rectified Attention を用いたエントロピー誘導型ビジュアルトークン解析フレームワーク ERA を提案する。
ERAはDEP(Dual-view Entropy Pruning)、BTR(Bias-Aware Tokencycle)、LAR(Logit-Reserving Attention Rectification)の3つの重要な構成要素から構成される。
論文 参考訳(メタデータ) (2026-06-30T17:20:29Z) - See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs [81.71384150441163]
大規模視覚言語モデル(LVLM)のためのコンテキスト認識注意介入(CAI)を提案する。
CAIは必要なときのみ2軸選択によってシーズを強制する。
複数のLVLMバックボーンとベンチマークの実験は、CAIが最先端の幻覚緩和を達成することを示している。
論文 参考訳(メタデータ) (2026-06-29T06:35:47Z) - Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory [24.777344301465064]
幻覚は人のような注意をそらす現象と強く結びついている。
本稿では,注目度向上による注意の注意散らしを補正する,改善されたイメージ知覚のための注意焦点付きアプローチを提案する。
論文 参考訳(メタデータ) (2026-05-23T14:36:06Z) - See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment [16.616065291567445]
MLLM(Multimodal large language model)は視覚入力を欠いたオブジェクトを幻覚させる。
DOP-OBCは、公平な注意の原則に基づいて構築された、トレーニング不要でアーキテクチャに依存しないデコーディング戦略である。
論文 参考訳(メタデータ) (2026-04-10T11:01:48Z) - SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation [33.381194425912234]
視覚言語モデル(VLM)はしばしば幻覚に悩まされ、視覚入力と矛盾するコンテンツを生成する。
SAGE, Sink-Aware Grounded Decoding frameworkは, 生成中の自己注意を動的に調節することで幻覚を緩和する。
本手法は,MSCOCOでは10.65%,AMBERでは7.19%の相対的改善を実現している。
論文 参考訳(メタデータ) (2026-03-29T22:52:03Z) - Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation [51.743225614196774]
マルチモーダル大言語モデル (MLLM) は視覚言語推論において顕著な進歩を遂げている。
彼らは幻覚に弱いままであり、そこで生成されたコンテンツは視覚的証拠から逸脱する。
近年の視覚強調法では、復号時に視覚トークンを補強することでこの問題に対処しようとしている。
本稿では,MLLMのトレーニングフリーフレームワークであるAdaptive Visual Reinforcement (AIR)を提案する。
論文 参考訳(メタデータ) (2026-02-27T14:18:51Z) - Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs [9.043999205886658]
大きな視覚言語モデルにおける幻覚は、言語が視覚的証拠を支配するときにしばしば起こる。
本稿では,視覚言語と言語のみの注意経路を構築するために,自己注意層内で動作するシングルパス機構であるContrastive Guidance(ACG)を提案する。
ACGは、計算コストを大幅に削減しつつ、最先端の忠実さとキャプション品質を達成する。
論文 参考訳(メタデータ) (2026-01-20T08:04:18Z) - Explaining multimodal LLMs via intra-modal token interactions [55.27436637894534]
MLLM(Multimodal Large Language Models)は、様々な視覚言語タスクにおいて顕著な成功を収めているが、その内部決定機構は十分に理解されていない。
モーダル内相互作用を利用した解釈可能性の向上を提案する。
論文 参考訳(メタデータ) (2025-09-26T14:39:13Z) - Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs [62.9348974370985]
約ゼロの余剰コストで幻覚を緩和するための注意再配置(AttnReal)を提案する。
我々のアプローチは,MLLMの注意分布が,歴史的出力トークンによって特徴が支配されるという重要な観測によって動機付けられている。
この観測に基づいて、AttnRealは出力トークンからの過剰な注意をリサイクルし、それを視覚トークンに再配置することで、MLLMの言語優先への依存を軽減します。
論文 参考訳(メタデータ) (2025-03-11T11:52:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。