論文の概要: Hidden not Deleted: How Networks Suppress Entangled Features
- arxiv url: http://arxiv.org/abs/2609.27593v1
- Date: Wed, 23 Sep 2026 09:10:12 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-25 00:05:17.965051
- Title: Hidden not Deleted: How Networks Suppress Entangled Features
- Title(参考訳): 隠された未削除:どうやってネットワークが絡み合った機能を抑圧するか
- Abstract要約: 線形射影を通した概念消去手法は,特徴が分離可能な部分空間を占めることを前提としている。
2つの特徴が1つのサブスペースを共有する反ポッド対に強制されると、最先端の線形消去はターゲットだけでなく両方を破壊する。
勾配勾配勾配で訓練されたネットワークは、この問題を非線形に解決するが、一様ではない。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.
- Abstract(参考訳): 線形射影を通して動作する概念消去法は、特徴が分離可能な部分空間を占めると仮定する。
2つの特徴が1つのサブスペースを共有する反ポッド対に強制されると、最先端の線形消去はターゲットだけでなく両方を破壊する。
勾配勾配勾配で訓練されたネットワークは、この問題を非線形に解くが、一様ではなく、鏡と影解と呼ばれる初期化に依存する2つの異なる回路レベルの解の1つに収束する。
この分岐を特徴の絡み合いの関数としてマッピングし、設定のアーティファクトではなく、安定したアトラクタ構造を反映していることを示し、両方のソリューションが、さらなるトレーニングを必要とせず、単一のスカラーパッチを通じて、消去された特徴の表現の実質的、測定可能なトレースを残していることを示すために、目的の因果介入を使用する。
このことは、最近LLMアンラーニングで実証的に観察された障害モードを反映しており、削除ではなく抑制によって忘れられた知識が再浮上する。
関連論文リスト
- Linear Adversarial Concept Erasure [108.37226654006153]
与えられた概念に対応する線形部分空間の同定と消去の問題を定式化する。
提案手法は, トラクタビリティと解釈性を維持しつつ, 深い非線形分類器のバイアスを効果的に軽減し, 高い表現性を有することを示す。
論文 参考訳(メタデータ) (2022-01-28T13:00:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。