論文の概要: Towards Robustness against Typographic Attack with Training-free Concept Localization
- arxiv url: http://arxiv.org/abs/2607.02494v1
- Date: Thu, 02 Jul 2026 17:55:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.960119
- Title: Towards Robustness against Typographic Attack with Training-free Concept Localization
- Title(参考訳): 学習自由概念ローカライゼーションによるタイポグラフィー攻撃に対するロバストネス
- Authors: Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang,
- Abstract要約: Contrastive Language-Image Pretraining Model (CLIP)は、現代のLVLM(Large Vision Language Models)の基盤となるビジョンエンコーダとして機能する。
CLIPモデルには重要な障害モードがあり、画像内に現れる無関係なテキストは、真の視覚的意味論ではなく、語彙的な意味に偏っている。
この問題は、一般的にはTypographic Attack (TA)と呼ばれ、自律運転のような安全クリティカルなアプリケーションに重大なリスクをもたらす脆弱性を露呈する。
そこで本稿では,TAに対する解釈可能かつ効果的なロバスト性を実現するための,新しい学習自由な機械的解釈法を提案する。
- 参考スコア(独自算出の注目度): 42.575178574443946
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve interpretable and effective robustness against TA, we propose a novel, training-free mechanistic interpretability method. Our method provides sampling-based interpretations of hidden state representations and quantitatively attributes semantic versus lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, we isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, thereby identifying the mechanistic source of TA. We further show that simple interventions applied directly to the identified circuits, without any additional training, can substantially improve robustness against Typographic Attacks in object classification. These interventions, such as selective adjustment of attention weights, also outperform both supervised and training-free defense methods. Our experiments demonstrate that applying the proposed intervention to the vision encoders of several state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under Typographic Attack interference on RIO-Bench. These results confirm both the efficacy and the generalizability of our mechanistic approach. Code is released at https://github.com/Liu-524/SamplingTAR.
- Abstract(参考訳): Contrastive Language-Image Pretraining (CLIP) を通じて訓練されたモデルは、現代のLVLM(Large Vision Language Models)の基盤となるビジョンエンコーダとして機能する。
画像内に現れる無関係なテキストは、真の視覚的意味論ではなく、語彙的な意味に偏っている。
タイポグラフィー攻撃(TA)と呼ばれるこの堅牢性問題は、自律運転のような安全クリティカルなアプリケーションに重大なリスクをもたらす脆弱性を露呈する。
TAに対する解釈可能かつ効果的なロバスト性を実現するために,新しい学習自由な機械的解釈法を提案する。
本手法は,隠れ状態表現のサンプリングに基づく解釈と,個々の注意対象に対する語彙的焦点に対する意味的意味の定量化を提供する。
確率論的解析と回路マイニングにより、語彙情報を不均等に符号化する特定のビジョントランス (ViT) 成分を分離し、TAの力学源を同定する。
さらに,物体分類におけるタイポグラフィー攻撃に対するロバスト性を大幅に向上させることができることを示す。
これらの介入、例えば注意重みの選択的な調整は、監督された防御法と訓練なし防御法の両方を上回ります。
提案手法を複数の最先端LVLMの視覚エンコーダに適用することにより,RIO-Benchに対するタイポグラフィー攻撃による視覚質問応答精度が大幅に向上することを示した。
これらの結果は,機械的アプローチの有効性と一般化性の両方を裏付けるものである。
コードはhttps://github.com/Liu-524/SamplingTARで公開されている。
関連論文リスト
- Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models [17.259725776748482]
頑健な微調整のための既存の敵の訓練手法は、視覚的堅牢性を高める上での言語の役割を概ね見落としている。
本研究では,QT-AFT(Quality Text-guided Adversarial Fine-Tuning)を提案する。
QT-AFTは、16のゼロショットデータセットで評価された、最先端のゼロショット対向ロバスト性とクリーンな精度を達成する。
論文 参考訳(メタデータ) (2025-07-22T06:13:30Z) - Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion [17.18411620606476]
本稿では,歩行者画像の微細な意味的特徴を損なうために,Attribute-aware Prompt Attack (AP-Attack)を導入する。
AP-Attackは最先端の転送可能性を実現し、従来の手法よりも22.9%上回った。
論文 参考訳(メタデータ) (2025-02-27T02:32:58Z) - Adversarial Robustification via Text-to-Image Diffusion Models [56.37291240867549]
アドリラルロバスト性は、ニューラルネットワークをエンコードする難しい性質として伝統的に信じられてきた。
データを使わずに敵の堅牢性を実現するために,スケーラブルでモデルに依存しないソリューションを開発した。
論文 参考訳(メタデータ) (2024-07-26T10:49:14Z) - MirrorCheck: Efficient Adversarial Defense for Vision-Language Models [55.73581212134293]
本稿では,視覚言語モデルにおける対角的サンプル検出のための,新しい,しかしエレガントなアプローチを提案する。
本手法は,テキスト・トゥ・イメージ(T2I)モデルを用いて,ターゲットVLMが生成したキャプションに基づいて画像を生成する。
異なるデータセットで実施した経験的評価により,本手法の有効性が検証された。
論文 参考訳(メタデータ) (2024-06-13T15:55:04Z) - SA-Attack: Improving Adversarial Transferability of Vision-Language
Pre-training Models via Self-Augmentation [56.622250514119294]
ホワイトボックスの敵攻撃とは対照的に、転送攻撃は現実世界のシナリオをより反映している。
本稿では,SA-Attackと呼ばれる自己拡張型転送攻撃手法を提案する。
論文 参考訳(メタデータ) (2023-12-08T09:08:50Z) - Proactive Pseudo-Intervention: Causally Informed Contrastive Learning
For Interpretable Vision Models [103.64435911083432]
PPI(Proactive Pseudo-Intervention)と呼ばれる新しい対照的な学習戦略を提案する。
PPIは、因果関係のない画像の特徴を保護するために積極的に介入する。
また,重要な画像画素を識別するための,因果的に通知された新たなサリエンスマッピングモジュールを考案し,モデル解釈の容易性を示す。
論文 参考訳(メタデータ) (2020-12-06T20:30:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。