論文の概要: TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
- arxiv url: http://arxiv.org/abs/2610.06994v1
- Date: Sun, 04 Oct 2026 07:42:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 02:58:29.498565
- Title: TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
- Title(参考訳): タレ:バックドアディフェンスの値段を読む前に、無害の双子に会いたい
- Abstract要約: バックドアディフェンスのリーダーボードは、クリーンな精度の低下を印刷し、除去コストとして読み取る。
毒を盛った被害者だけで測定すると、この落下はどんなモデルでも防御効果から切り離すことはできない。
我々は,3キーパッチ,署名付きタレコラム,およびシード安定防御のためのツインフリー推定器TARE-Zを出荷する。
- 参考スコア(独自算出の注目度): 6.602613110804935
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR that never fires, so their victims never anneal and are the least accurate in 30/31 public CIFAR cells at $\leq$5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to their own level, so there the published cost is negative, the benchmark's rating clips the "gain" to zero, and 2 of 48 citing defense papers we read rest a no-cost claim on those cells; TSBD and CGD, re-run with their code, "gain" on a never-poisoned model too. A $2\times2$ editing only that scheduler line isolates the cause, its swapped arms self-registered before they ran: the sign of the fine-tuning family's clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. TARE runs the same defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, $\leq$10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On the BadNets grid seven of eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 (fine-tuning) to +5.70 points (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share unidentified. We ship the three-key patch, a signed tare column (7 attacks $\times$ 8 defenses) and TARE-Z, a twin-free estimator for seed-stable defenses.
- Abstract(参考訳): バックドアディフェンスのリーダーボードは、クリーンな精度の低下を印刷し、除去コストとして読み取る。
BackdoorBenchの16件の攻撃のうち3件は、設定ファイルである: WaNet, BPP and Input-Aware ship a MultiStepLR that never fire, so their victims never anneal, and are least accurate in 30/31 public CIFAR cells at $\leq$5%。
PreAct-ResNet18では、微調整されたファミリーディフェンスが低いスタートを自身のレベルに戻すため、公開コストは負であり、ベンチマークのレーティングは"ゲイン"をゼロにし、48のうち2つは、私たちが読み取った防衛文書を引用するものであり、TSBDとCGDはコードで再実行される。
細調整された家族のクリーンモデルコストのサインは両面を逆転させ、熱傷を受けた被害者の利益は、スケジュールの後に0、44/44の種だけ減少し、BPP、FT-SAM、CIFAR-100、VGG19-BNで複製され、第2のツールキットで誘導される。
TAREは、同じレシピ、スケジュール、シード(BackdoorBench、$\leq$10の有毒な画像、5%の攻撃成功しか認められていない)の無害双生児に対して同じ防御を行う。
BadNetsグリッドの8つの防御線のうち7つは双子(ニューラル・クリーズ)を充電し、+0.13(微調整)を+5.70ポイント(I-BAU)とし、8番目のABLはそれを破壊する。
攻撃中、スタートはランキングからキャンセルされるので、タレはそこに何の注文もせず、2つの推定値の下に追加される毒物は修正されず、除去のシェアは未確認である。
私たちは3キーパッチ、署名付きタレコラム(7つの攻撃で$\times$8ディフェンス)、およびシード安定ディフェンスのためのツインフリーな推定器TARE-Zを出荷します。
関連論文リスト
- Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks [52.041551928780244]
バックドアの毒殺攻撃は、さもなければクリーンな微調整データに有毒な例を付加する。
これは最悪のケースの脆弱性を過小評価する可能性がある。
SAILS (Set-level Audit-Informed Iterative Learned Selection) を導入する。
論文 参考訳(メタデータ) (2026-09-14T04:44:28Z) - RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning [0.0815557531820863]
本稿では,パイプラインに対する擬似コーパス攻撃に対する層状防御であるRAGuardを紹介する。
第1層は、合成有毒文書(偽事実、矛盾、推論トラップ)に密集したレトリバーを微調整する。
2番目の層であるZeroKnowledge Patch ZKIPは、ラベルのないブラックボックスフィルタである。
論文 参考訳(メタデータ) (2026-07-28T23:21:04Z) - Cordyceps: Covert Control Attacks on LLMs via Data Poisoning [7.104161390387404]
大規模言語モデル(LLM)は、しばしば敵に毒を盛る未処理のテキストデータセットに基づいて微調整される。
本稿では, LLM に情報隠蔽方式を確実に, ひそかに教えるデータ中毒手法を提案する。
引き起こされた隠蔽スキームは、任意の悪意のある命令をエンコードし、デコードするので、新しく微妙な毒によって引き起こされる脆弱性が明らかになる。
論文 参考訳(メタデータ) (2026-05-26T06:28:30Z) - Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget [23.409081157517264]
Checkerboardは理論的に根拠のない、学習のないクリーンラベルのバックドア攻撃だ。
線形分離性定式化から閉形式にトリガを導出し,サロゲートモデルトレーニングとトリガ最適化の必要性を除去する。
4つのベンチマークデータセットで、Checkerboardは8つのベースラインアタックを上回り、毒の少ない予算下で最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2026-05-02T07:14:49Z) - P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs [49.908234151374785]
微調整の間、大規模言語モデル(LLM)は、データポゾンによるバックドア攻撃に対してますます脆弱である。
汎用的で効果的なバックドアディフェンスアルゴリズムであるPoison-to-Poison (P2P)を提案する。
P2Pはタスク性能を維持しながら悪質なバックドアを中和できることを示す。
論文 参考訳(メタデータ) (2025-10-06T05:45:23Z) - Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising [56.04951180983087]
我々は、$ell$-normの条件で、クリーンラベル毒殺攻撃に対する認証された防御を提示する。
$randomized$$smoothingによって達成された対向的堅牢性に触発されて、オフザシェルフ拡散復調モデルが、改ざんしたトレーニングデータの衛生化をいかに行うかを示す。
論文 参考訳(メタデータ) (2024-03-18T17:17:07Z) - Adversarial Attack on Attackers: Post-Process to Mitigate Black-Box
Score-Based Query Attacks [25.053383672515697]
本稿では,攻撃者に対する敵攻撃(AAA)という新たな防御策を提案し,SQAを誤った攻撃方向へ誘導する。
このように、SQAはモデルの最悪のケースの堅牢性に関係なく防止される。
論文 参考訳(メタデータ) (2022-05-24T15:10:50Z) - A Partial Break of the Honeypots Defense to Catch Adversarial Attacks [57.572998144258705]
検出正の正の率を0%、検出AUCを0.02に下げることで、この防御のベースラインバージョンを破る。
さらなる研究を支援するため、攻撃プロセスの2.5時間のキーストローク・バイ・キーストロークスクリーンをhttps://nicholas.carlini.com/code/ccs_honeypot_breakで公開しています。
論文 参考訳(メタデータ) (2020-09-23T07:36:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。