論文の概要: Backdoor as Probe: Test-Time Adversarial Defense for CLIP
- arxiv url: http://arxiv.org/abs/2609.34641v1
- Date: Mon, 28 Sep 2026 08:53:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 11:39:01.954085
- Title: Backdoor as Probe: Test-Time Adversarial Defense for CLIP
- Title(参考訳): バックドア・アズ・プローブ:CLIPのテストタイム・アドバイザリ・ディフェンス
- Abstract要約: テストタイムの敵防衛は、CLIPのようなビジョン言語基盤モデルの堅牢性を改善する。
バックドアのトリガー・ツー・ターゲット機構を再利用することにより、敵のアクティベーションシフトを防御信号に変換する。
この知見に基づいて,CLIP に対するテスト時対抗防御である Probe (BaP) として emphBackdoor を提案する。
- 参考スコア(独自算出の注目度): 64.795947088543
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose \emph{Backdoor as Probe} (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models. Project page: https://robin-wzq.github.io/Backdoor-as-Probe/
- Abstract(参考訳): テストタイムの敵防衛は、CLIPのようなビジョン言語基盤モデルの堅牢性を改善する。
しかしながら、敵のアクティベーションシフトは、通常、活用する信号ではなく、抑制する歪みとして扱われる。
バックドアのトリガー・ツー・ターゲット機構を再利用することで,これらのシフトを防御信号に変換する。
鍵となるのは、ディフェンダーが制御するバックドアを、クリーンな入力によって弱活性化されるが、敵のシフトによって強く活性化されるプローブとして埋め込むことである。
この知見に基づいて,CLIP に対するテスト時対角防御である emph{Backdoor as Probe} (BaP) を提案する。
BaPは、選択したMLP層へのクローズドフォームモデル編集を通じてプローブを構成する。
平均対向活性化シフトとディフェンダー特定意味方向を各層の低エネルギー入力と出力活性化部分空間に投影し、それぞれトリガーと目標方向を得る。
推定時、敵入力は目標方向に沿って測定可能な応答を生成して検出する。
その後、BaPは検出された入力を選択的に修正し、小さな摂動を最適化し、その表現を敵のシフトから切り離し、クリーンな部分空間へと誘導する。
16のベンチマークでの実験では、BaPは平均ロバスト精度を1.0\%から52.3\%に改善し、精度を保ちながら、(5.7\times\)推論スピードアップまでの最先端の手法に匹敵する性能を達成している。
BaPはさらに、大きな視覚言語モデルに対する敵攻撃の一般化を示す。
プロジェクトページ:https://robin-wzq.github.io/Backdoor-as-Probe/
関連論文リスト
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models [17.839413035304748]
LLM(Large Language Models)に対するバックドアのアンアライメント攻撃は、隠れたトリガーを使用して、安全アライメントのステルスな妥協を可能にする。
我々は,裏口LDMを不活性化させるために,推論中にトリガサンプルを検出するブラックボックスディフェンスBEATを紹介する。
本手法は, サンプル依存目標の課題を, 反対の観点から解決する。
論文 参考訳(メタデータ) (2025-06-19T16:30:56Z) - Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models [70.03122709795122]
テキストと画像の拡散モデルをターゲットにしたバックドア攻撃が急速に進んでいる。
現在のバックドアサンプルは良性サンプルと比較して2つの重要な異常を示すことが多い。
我々はこれらの成分を明示的に緩和することでTwT(Trigger without Trace)を提案する。
論文 参考訳(メタデータ) (2025-03-22T10:41:46Z) - Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP [51.04452017089568]
CBPT(Class-wise Backdoor Prompt Tuning)は、テキストプロンプトでCLIPを間接的に浄化する効率的な防御機構である。
CBPTは、モデルユーティリティを保持しながら、バックドアの脅威を著しく軽減する。
論文 参考訳(メタデータ) (2025-02-26T16:25:15Z) - T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models [70.03122709795122]
バックドア攻撃の検出, 局所化, 緩和のための総合防御手法T2IShieldを提案する。
バックドアトリガーによって引き起こされた横断アテンションマップの「アシミレーション現象」を見いだす。
バックドアサンプル検出のために、T2IShieldは計算コストの低い88.9$%のF1スコアを達成している。
論文 参考訳(メタデータ) (2024-07-05T01:53:21Z) - LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors [10.136109501389168]
LMSanitatorは、Transformerモデル上でタスク非依存のバックドアを検出し、削除するための新しいアプローチである。
LMSanitatorは960モデルで92.8%のバックドア検出精度を達成し、ほとんどのシナリオで攻撃成功率を1%以下に下げる。
論文 参考訳(メタデータ) (2023-08-26T15:21:47Z) - Backdoor Mitigation by Correcting the Distribution of Neural Activations [30.554700057079867]
バックドア(トロイジャン)攻撃はディープニューラルネットワーク(DNN)に対する敵対的攻撃の重要なタイプである
バックドア攻撃の重要な特性を解析し、バックドア・トリガー・インスタンスの内部層活性化の分布の変化を引き起こす。
本稿では,分散変化を補正し,学習後のバックドア緩和を効果的かつ効果的に行う方法を提案する。
論文 参考訳(メタデータ) (2023-08-18T22:52:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。