論文の概要: Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning
- arxiv url: http://arxiv.org/abs/2608.04347v1
- Date: Wed, 05 Aug 2026 01:47:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.685628
- Title: Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning
- Title(参考訳): 鏡をみる:微調整による副作用の検証
- Authors: Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto, Yuji Naraki, Ryotaro Shimizu, Wenya Wang,
- Abstract要約: 微調整により、ソースモデルはターゲットドメインで望ましい機能や振る舞いを取得できる。
微調整は、ソースモデルに存在したアライメント特性を劣化させることもできる。
イントロスペクションアダプタは微調整によって引き起こされる行動変化を記述する。
- 参考スコア(独自算出の注目度): 17.810096531720927
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can be trained using LoRA-based modules known as introspection adapters (IAs) to describe behavioral changes induced by fine-tuning. However, existing studies primarily consider settings in which the model is fine-tuned on datasets explicitly designed to implant a specific behavior and is then asked to explain the implanted behavior. This differs from practical deployment scenarios, where the central concern is often side-effect misalignment: unintended degradation of alignment caused by fine-tuning on tasks that are not obviously related to safety or alignment. To bridge this gap, we formulate a novel problem setting called \emph{side-effect introspection}, in which the target of introspection is not a behavior explicitly implanted through fine-tuning, but rather alignment shifts that emerge as unintended side effects, and we construct a dataset for this setting. Furthermore, to enhance sensitivity to internal model changes, we propose the Delta-Aware Introspection Adapter (DAIA), a novel mechanism designed to explicitly process both base-model activations and activation differences induced by fine-tuning. Our empirical evaluation shows that introspection learning generalizes to unseen fine-tuned models and safety categories, and that DAIA consistently outperforms existing introspection adapters.
- Abstract(参考訳): ファインチューニングにより、ソースモデルは、その汎用能力の多くを保持しながら、ターゲットドメインで望ましい能力と振舞いを取得することができる。
しかし、この適応プロセスは、ソースモデルに存在するアライメント特性を劣化させることもできる。
近年の研究では、イントロスペクションアダプタ(IA)として知られるLoRAベースのモジュールを使用して、大規模な言語モデルをトレーニングすることで、微調整によって引き起こされる振る舞いの変化を記述できることが示されている。
しかし、既存の研究では、モデルを特定の振る舞いを埋め込むように明示的に設計されたデータセットに基づいて微調整し、埋め込みされた振る舞いを説明するように要求される設定を主に検討している。
安全性やアライメントと明らかに関係のないタスクの微調整によるアライメントの意図しない劣化です。
このギャップを埋めるために、我々は「emph{side-effect Introspection」と呼ばれる新しい問題設定を定式化し、イントロスペクションの対象は、微調整によって明示的に埋め込まれた振る舞いではなく、意図しない副作用として現れるアライメントシフトであり、この設定のためのデータセットを構築する。
さらに、内部モデル変更に対する感度を高めるために、ベースモデルアクティベーションと微調整によるアクティベーション差の両方を明示的に処理する新しいメカニズムである、Delta-Aware Introspection Adapter (DAIA)を提案する。
我々の経験的評価は、イントロスペクション学習が未確認の微調整モデルや安全カテゴリーに一般化し、DAIAが既存のイントロスペクションアダプタより一貫して優れていることを示している。
関連論文リスト
- Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation [74.17379276939599]
近年のQwen-3.5シリーズにおいても,アクティベーションステアリングが広範囲のアライメントを引き起こすことが示されている。
ステアリングサイズ, ステアリングサブスペースの低ランク構造, ステアリングベクター構築時のエポック数など, キーステアリング固有の因子を解析することにより, AS誘起EMの特性を特徴づける。
論文 参考訳(メタデータ) (2026-06-07T15:34:59Z) - Visual prompting reimagined: The power of the Activation Prompts [72.85146015928626]
本稿では,入力レベルのVPの範囲を広げる,アクティベーションプロンプト(AP)の概念を導入する。
APは畳み込みニューラルネットワークと視覚変換器の正規化チューニングと密接に関連している。
畳み込みニューラルネットワークと視覚変換器の正規化チューニングとAPは密接に関連していることを示す。
論文 参考訳(メタデータ) (2026-04-07T20:28:24Z) - Alignment-Aware Model Adaptation via Feedback-Guided Optimization [27.93864970404945]
ファインチューニングは、ファンデーションモデルを下流タスクに適応するための主要なメカニズムである。
本稿では,外部アライメント信号からのフィードバックをポリシー段階の正規化を通じて統合するアライメント対応微調整フレームワークを提案する。
論文 参考訳(メタデータ) (2026-02-02T16:03:16Z) - Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures [70.48661957773449]
創発的ミスアライメント(英: Emergent Misalignment)とは、狭い範囲のデータに対する微調整された大きな言語モデルによって、広範囲に不整合な振る舞いが引き起こされる障害モードを指す。
複数のドメインやモデルファミリにまたがって、特定の文字レベルの配置を示すデータの微調整モデルは、誤操作よりもはるかに強く、転送可能な微調整を誘導する。
論文 参考訳(メタデータ) (2026-01-30T15:28:42Z) - When Domain Pretraining Interferes with Instruction Alignment: An Empirical Study of Adapter Merging in Medical LLMs [0.6345523830122167]
大規模言語モデルは、ドメイン適応と命令アライメントを組み合わせる際に驚くべきアダプタ干渉を示す。
医学LLMのための2段階のLORAパイプラインについて検討し、ドメイン指向事前トレーニング(PT)と教師付き微調整(SFT)を個別に訓練し、後にマージした。
論文 参考訳(メタデータ) (2026-01-26T10:54:06Z) - Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs [0.0]
安全でないコードに対する微調整は、アライメントに反する内部的な変更を誘発することを示す。
我々は、アライメントの振る舞いを管理するモデルの活性化空間における共有潜在次元を同定する。
論文 参考訳(メタデータ) (2025-07-04T15:36:58Z) - Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction [57.19302613163439]
モデル適応のための統一フレームワークとして,ニューラルネットワークの再プログラム可能性を導入する。
本稿では,4つの重要な側面にまたがる情報操作アプローチを分類する分類法を提案する。
残る技術的課題や倫理的考察も分析する。
論文 参考訳(メタデータ) (2025-06-05T05:42:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。