論文の概要: PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders
- arxiv url: http://arxiv.org/abs/2608.10985v1
- Date: Tue, 11 Aug 2026 14:38:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-12 19:14:46.07764
- Title: PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders
- Title(参考訳): PEAK:k-スパースオートエンコーダによる精密で永続的な概念消去
- Authors: Man Jiang, Ouxiang Li, Weibao Xue, Zhenhua Tang, Yuan Wang, Shuo Wang, Yanbin Hao,
- Abstract要約: 大規模なテキスト・ツー・イメージ(T2I)拡散モデルによる概念の消去は、著作権侵害、プライバシー侵害、攻撃的コンテンツに対する懸念が高まっているため、ますます重要になっている。
我々は,k-Sparse Autoencoders (kSAEs) を用いたtextbftextitprecise と textbftextitpersistent の概念消去フレームワーク PEAK を提案する。
- 参考スコア(独自算出の注目度): 25.995968388013555
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, privacy violations, and offensive content. Existing approaches struggle to achieve both precise and persistent concept erasure: inaccurate localization of concept-related representations may cause unintended semantic interference, while incomplete removal of the underlying concept knowledge allows adversarial recovery. To address this dilemma, we propose PEAK, a \textbf{\textit{precise}} and \textbf{\textit{persistent}} concept erasure framework via k-Sparse Autoencoders (kSAEs). PEAK first trains a kSAE on internal activations of the diffusion denoising network to decompose dense representations into interpretable sparse features. By contrasting sparse activations induced by target and non-target prompts, PEAK identifies a compact set of target-specific features according to both activation strength and frequency. These localized features are then used for parameter optimization, where PEAK selectively suppresses target-related activations while preserving complementary non-target ones towards the original model. This feature-guided optimization embeds concept erasure directly into diffusion parameters, eliminating the need for additional inference-time intervention and facilitating effective persistence against adversarial attacks. Extensive experiments demonstrate that PEAK achieves effective and robust concept erasure. On the I2P benchmark, PEAK reduces NudeNet detections from 582 to 6, lowers the average attack success rate (ASR) from 96.52\% to 5.63\%, and preserves general generation quality on MS-COCO with a near-zero KID. Our code and models are available at: https://github.com/manmanTAT/PEAK
- Abstract(参考訳): 大規模なテキスト・ツー・イメージ(T2I)拡散モデルによる概念の消去は、著作権侵害、プライバシー侵害、攻撃的コンテンツに対する懸念が高まっているため、ますます重要になっている。
既存のアプローチは、正確な概念消去と永続的な概念消去の両方を達成するのに苦労している: 概念関連表現の不正確な局所化は意図しない意味的干渉を引き起こすが、根底にある概念知識の不完全な除去は敵の回復を可能にする。
このジレンマに対処するため,k-Sparse Autoencoders (kSAEs) を介して, PEAK, a \textbf{\textit{precise}} と \textbf{\textit{persistent}} の概念消去フレームワークを提案する。
PEAKはまず、拡散分解ネットワークの内部活性化についてkSAEを訓練し、密度表現を解釈可能なスパース特徴に分解する。
PEAKは、ターゲットプロンプトと非ターゲットプロンプトによって誘導されるスパースアクティベーションと対照的に、アクティベーション強度と周波数の両方に応じて、ターゲット固有の特徴のコンパクトなセットを特定する。
これらの局所化機能はパラメータ最適化に使用され、PEAKは元のモデルに相補的な非ターゲットを保存しながら、ターゲット関連のアクティベーションを選択的に抑制する。
この特徴誘導最適化は、概念消去を直接拡散パラメータに埋め込んで、追加の推論時間介入の必要性を排除し、敵攻撃に対する効果的な永続性を促進する。
大規模な実験は、PEAKが効果的で堅牢な概念消去を達成することを示した。
I2Pベンチマークでは、PEAKはNudeNet検出を582から6に減らし、平均攻撃成功率(ASR)を96.52\%から5.63\%に下げ、ほぼゼロのKIDでMS-COCOの一般的な生成品質を維持する。
私たちのコードとモデルは、https://github.com/manmanTAT/PEAKで利用可能です。
関連論文リスト
- Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization [51.33055188976282]
敵対的攻撃は 個人化された画像生成において 悪意のあるコンテンツの偽造を防ぐ 防衛戦略として浮上した
本稿では,2段階遅延特徴最適化 (TS-LFO) を導入する。
広範な実験により、TS-LFOは、最先端のSOTA著作権防衛を一貫して回避し、DiffPure、GrIDPure、IMPRESSなどのSOTA著作権攻撃より優れていることが示されている。
論文 参考訳(メタデータ) (2026-06-06T07:59:08Z) - DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models [55.30555646945055]
テキスト・ツー・イメージ(T2I)モデルはセマンティック・リークに対して脆弱である。
DeLeakerは、モデルのアテンションマップに直接介入することで、漏洩を緩和する軽量なアプローチである。
SLIMはセマンティックリークに特化した最初のデータセットである。
論文 参考訳(メタデータ) (2025-10-16T17:39:21Z) - CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models [7.68494752148263]
CUREは、事前訓練された拡散モデルの重み空間で直接動作する、トレーニング不要の概念未学習フレームワークである。
スペクトル消去器は、安全な属性を保持しながら、望ましくない概念に特有の特徴を特定し、分離する。
CUREは、対象とする芸術スタイル、オブジェクト、アイデンティティ、明示的なコンテンツに対して、より効率的で徹底的な除去を実現する。
論文 参考訳(メタデータ) (2025-05-19T03:53:06Z) - SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models [56.83154571623655]
モデルパラメータを直接編集する効率的な概念消去手法であるSPEEDを導入する。
Speedyは、パラメータ更新がターゲット以外の概念に影響しないモデル編集スペースであるnullスペースを検索する。
たった5秒で100のコンセプトを消去しました。
論文 参考訳(メタデータ) (2025-03-10T14:40:01Z) - Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models [76.39651111467832]
本稿では,Reliable and Efficient Concept Erasure (RECE)を提案する。
派生した埋め込みによって表現される不適切なコンテンツを緩和するために、RECEはそれらをクロスアテンション層における無害な概念と整合させる。
新たな表現埋め込みの導出と消去を反復的に行い、不適切な概念の徹底的な消去を実現する。
論文 参考訳(メタデータ) (2024-07-17T08:04:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。