論文の概要: Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning
- arxiv url: http://arxiv.org/abs/2608.24482v1
- Date: Tue, 25 Aug 2026 12:28:55 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:34.967566
- Title: Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning
- Title(参考訳): 静的解釈性を超えて:より良いチューニングのためのSFT前パラメータからのSFT後メカニズムの予測
- Authors: Hang Chen, Jiaying Zhu, Wenya Wang,
- Abstract要約: 機械的ローカライゼーションは機械的解釈可能性と後学習最適化を橋渡しする。
本研究では,SFT後の解釈可能性状態を正確に推定するフォワード・ルックティング・フレームワークを提案する。
- 参考スコア(独自算出の注目度): 28.24632224083753
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative approaches and then guiding parameter-efficient Supervised Fine-Tuning (SFT) in a ``locating-then-tuning'' paradigm. However, due to the retrospective nature of mechanistic interpretability, directly interpreting pre-SFT models introduces misleading conclusions. Specifically for novel tasks, initially identified neurons differ drastically from those governing the final model, introducing biases that actively disrupt SFT. To address this, we propose a forward-looking localization framework that accurately estimates the post-SFT interpretability state using only pre-SFT parameters and the target dataset. Theoretically, we model SFT as a continuous parameter evolution, leveraging Taylor expansion to rigorously bridge the post-tuning mechanistic objective with the pre-SFT model's dynamic gradients. Practically, we design dual-granularity (neuron- and component-level) localization pipelines. Extensive experiments demonstrate that our approach not only provides superior SFT guidance but also exhibits robust performance and temporal scalability across increasing model sizes. This work transcends the fundamental limitation of traditional interpretability-its inability to identify task-critical mechanisms before they are trained-pioneering a predictive frontier that unites mechanistic interpretability with targeted optimization.
- Abstract(参考訳): メカニカルローカライゼーションは、解釈的アプローチを通じて臨界パラメータを分離し、次にパラメータ効率の高いスーパーバイザード・ファイン・チューニング(SFT)を‘ロケーション・then-tuning’パラダイムで導くことによって、機械的解釈可能性と後学習の最適化を橋渡しする。
しかし、機械論的解釈可能性の振り返りの性質のため、SFT以前のモデルを直接解釈することは誤解を招く結論をもたらす。
具体的には、新しいタスクにおいて、最初に特定されたニューロンは最終モデルを決定するものとは大きく異なり、SFTを積極的に破壊するバイアスが導入された。
そこで本研究では,事前SFTパラメータとターゲットデータセットのみを用いて,SFT後の解釈可能性状態を正確に推定する,前方方向のローカライゼーションフレームワークを提案する。
理論的には、SFTを連続パラメータの進化としてモデル化し、テイラー展開を利用して、学習後の機械的目的をSFTモデルの動的勾配に厳密に橋渡しする。
実際、我々は二粒度(ニューロンレベルとコンポーネントレベルの)ローカライゼーションパイプラインを設計する。
大規模な実験により、我々のアプローチは優れたSFTガイダンスを提供するだけでなく、モデルのサイズが大きくなるにつれて、堅牢な性能と時間的スケーラビリティを示すことが示された。
この研究は、伝統的な解釈可能性の欠如の基本的な限界を超越し、目標とする最適化と機械的解釈可能性を統合する予測的フロンティアを訓練する前にタスククリティカルなメカニズムを識別する。
関連論文リスト
- LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure [16.743783987674842]
Supervised Fine-tuning (SFT) は、事前訓練された言語モデルを下流ドメインに適応するための標準的なアプローチである。
我々は,この固有エントロピー構造を明示的に保護するために,局所保存型スーパーバイザード・ファイン・チューニングの目的であるLP-SFTを提案する。
論文 参考訳(メタデータ) (2026-07-06T07:14:22Z) - PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective [52.693471818837395]
PEFTは安定性・塑性ジレンマにより評価されるべきである。
本稿では,下流性能と一般能力の維持を計測するベンチマークPEFT-Arenaを紹介する。
そこで本研究では,パスワイド巻き戻しによるポストホック改善の事例研究を行った。
論文 参考訳(メタデータ) (2026-05-27T17:59:51Z) - Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport [51.56484100374058]
デプロイ後、ユーザの好みが進化し、初期設定が望ましくないようになる。
我々は、観測データから、NNの条件付き出力分布がハイパーパラメータでどのように変化するかを学ぶ。
我々は、NNを観測されていないハイパーパラメータで近似する代理モデルを構築した。
論文 参考訳(メタデータ) (2026-03-02T11:55:02Z) - Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections [65.36449542323277]
本稿では,Large Language Model (LLM) 後の学習において,SFT(Supervised Fine-Tuning) と優先学習を統合した理論フレームワークを提案する。
そこで本研究では,学習率の簡易かつ効果的な削減手法を提案する。
論文 参考訳(メタデータ) (2025-06-15T05:42:29Z) - Forecast-PEFT: Parameter-Efficient Fine-Tuning for Pre-trained Motion Forecasting Models [68.23649978697027]
Forecast-PEFTは、モデルのパラメータの大部分を凍結し、新しく導入されたプロンプトとアダプタの調整に集中する微調整戦略である。
実験の結果,Forecast-PEFTは動作予測タスクにおいて従来のフルチューニング手法よりも優れていた。
Forecast-FTは予測性能をさらに改善し、従来のベースライン法よりも最大9.6%向上した。
論文 参考訳(メタデータ) (2024-07-28T19:18:59Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。