論文の概要: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
- arxiv url: http://arxiv.org/abs/2607.00572v3
- Date: Wed, 08 Jul 2026 05:02:03 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-09 16:11:06.584318
- Title: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
- Title(参考訳): HARC:ロバストな安全アライメントのためのハーモフルネスと拒絶方向の結合
- Authors: Shei Pern Chua, Hao Wu, Qianli Ma, Fangzhao Wu,
- Abstract要約: 本研究は,LLMのアライメントにより,残差ストリームにおける有害性と拒絶を,プロンプト側トークン位置における分離可能な方向としてエンコードしていることを示す。
本稿では,2方向をプロンプト位置と応答位置でペアリングする微調整手法であるHARCを紹介する。
- 参考スコア(独自算出の注目度): 23.776344382492997
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
- Abstract(参考訳): ジェイルブレイクが成功した理由を説明し、ロバストなアライメント戦略の設計を通知するからである。
先行研究では、LLMは有害性をエンコードし、プロンプト側トークン位置の残留ストリームにおいて分離可能な方向として拒絶することを示した。
本研究では, トークン発生前の拒絶方向, 有害方向のどちらかを抑えることで, 迅速な符号化に成功し, 有害面の分離可能な領域を占有する異なる攻撃クラスを有することを示す。
分析対象を対応点に拡張することで,入力が危険であると認識できなかった場合でも,その内容を生成しながら有害な内容を認識できることが判明した。
HARC (Harmfulness-And-Refusal Coupling) は2つの方向をプロンプト位置と応答位置でペアリングする微調整法である。
介入は有害な拒絶部分空間に制限されるため、残余のストリームの残りは無傷のまま残され、一般的な能力を低下したり、過剰な拒絶を減少させることはない。
広範な実験を通じて、HARCは、主要なトレーニング時間と推論時間安全手法にまたがる6つのベースラインの中で、強靭性-耐久性-使用性トレードオフを達成している。
5つのモデルファミリと2つのスケールにまたがるプロンプト位置と応答位置の有害さと拒絶方向をアーキテクチャ固有のチューニングなしで比較した。
関連論文リスト
- From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning [14.6508023458559]
既存の評価では、モデルが有害性を認識することを学んだかどうかを明らかにしていない。
本研究では, 有害性担体, 拒絶性担体, 結合性を測定する二重安全幾何プロトコルを用いて検討する。
論文 参考訳(メタデータ) (2026-06-15T07:50:00Z) - Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models [50.91504059485288]
本報告では,全頭部のグローバルな最適化により,安全クリティカルな注意点を同時に識別するフレームワークを提案する。
我々は,アクティベーション・リマッチによって同定された安全ベクトルを利用する,新しい推論時ホワイトボックス・ジェイルブレイク法を開発した。
論文 参考訳(メタデータ) (2026-01-22T09:32:43Z) - ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning [64.32925552574115]
ARMORは、jailbreak戦略を分析し、コアインテントを抽出する、大規模な言語モデルである。
ARMORは最先端の安全性能を達成し、平均有害率は0.002であり、高度な最適化ベースのジェイルブレイクに対する攻撃成功率は0.06である。
論文 参考訳(メタデータ) (2025-07-14T09:05:54Z) - Does Representation Intervention Really Identify Desired Concepts and Elicit Alignment? [73.80382983108997]
表現の介入(Representation intervention)は、大規模言語モデルにおいて基礎となる概念を符号化する表現の発見と修正を目的としている。
介入が忠実であれば、介入されたLLMは有害な概念を消去し、非分配的敵のプロンプトとアウト・オブ・ディストリビューションのジェイルブレイクの両方に対して堅牢であるべきである。
本研究では,有害表現と良性表現の境界を簡易化する概念集中(COCA)を提案する。
論文 参考訳(メタデータ) (2025-05-24T12:23:52Z) - Improving LLM Safety Alignment with Dual-Objective Optimization [81.98466438000086]
大規模言語モデル(LLM)の既存のトレーニング時間安全アライメント技術は、ジェイルブレイク攻撃に対して脆弱なままである。
本研究では,DPOの目的を2つの構成要素にまとめる安全アライメントの改善について提案する。(1) 安全でない世代が部分的に発生しても拒否を促す頑健な拒絶訓練,(2) 有害な知識の未学習。
論文 参考訳(メタデータ) (2025-03-05T18:01:05Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。