論文の概要: DemoPSD: Disagreement-Modulated Policy Self-Distillation
- arxiv url: http://arxiv.org/abs/2607.02502v2
- Date: Mon, 06 Jul 2026 11:02:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 17:33:48.995991
- Title: DemoPSD: Disagreement-Modulated Policy Self-Distillation
- Title(参考訳): DemoPSD: 分散政策による自己蒸留
- Authors: Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song,
- Abstract要約: 大規模言語モデルを訓練するための実践的な方法として、オンデマンド自己蒸留が登場している。
DemoPSDは、教師と生徒の分布の重み付けされた幾何学的組み合わせに向けて学生を操縦する。
- 参考スコア(独自算出の注目度): 49.30809363804017
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access. However, recent studies have found that the teacher's dense token-level supervision, conditioned on privileged information, can lead to overfitting to in-domain patterns, suppress exploration, and hurt cross-domain generalization, while also introducing a more fundamental issue: *privileged information leakage*, where the student encodes answer-dependent shortcuts that are unavailable at test time. We introduce **DemoPSD**, a novel framework that resolves such problems through the idea of *selective adoption of teacher guidance*. Instead of fitting the full teacher distribution, DemoPSD steers the student toward a *reverse-KL barycenter target*, a weighted geometric combination of the teacher and student distributions, that naturally balances learning from the teacher with preserving the student's own reasoning capacity. We measure the difference between their distributions and use such a discrepancy to adaptively control the blending at each token position. We provably show that DemoPSD achieves **(1)** *leakage attenuation*, i.e., effective mitigation of privileged information leakage; and **(2)** *exploration preservation*, i.e., preservation of exploration capacity under dense token-level distillation. Extensive experiments on SciKnowEval across four scientific fields show that DemoPSD outperforms both GRPO and SDPO while maintaining higher training entropy and robustly generalizing to out-of-distribution GPQA benchmarks.
- Abstract(参考訳): On-policy Self-distillation (OPSD) は、大きな言語モデル(LLM)を学習し、単一のモデルが教師と学生の両方に異なるレベルの情報アクセスを提供するための実践的な方法として登場した。
しかし、近年の研究では、特権情報に基づく教師の密集したトークンレベルの監督は、ドメイン内のパターンに過度に適合し、探索を抑え、クロスドメインの一般化を損なうとともに、より根本的な問題を提起している。
そこで我々は,**DemoPSD*を紹介した。**DemoPSD**は,教師指導の選抜的導入というアイデアを通じて,このような問題を解決する新しいフレームワークである。
DemoPSDは教師の分布を完全に調整する代わりに、教師と生徒の分布の重み付けされた幾何学的組み合わせである*リバース-KLのバリセンターターゲット*に向けて学生を操り、教師からの学習と生徒自身の推論能力のバランスを取る。
それらの分布の差を測定し,その差を利用して各トークン位置でのブレンディングを適応的に制御する。
我々は,DemoPSDが**(1)** *leakage attenuation*,すなわち,特権情報漏洩の効果的な緩和,**(2)***探索保存*,すなわち,密集したトークンレベルの蒸留条件下での探索能力の保存を実現することを確実に示す。
4つの科学分野にわたるSciKnowEvalの大規模な実験により、DemoPSDはGRPOとSDPOの両方に優れ、より高いトレーニングエントロピーを維持し、GPQAベンチマークに頑健に一般化していることが示された。
関連論文リスト
- DOPD: Dual On-policy Distillation [59.65986569617147]
オンライン蒸留は、高密度のトークンレベルの信号で学生サンプルの軌跡を監督することで、優れた容量転送を提供する。
特権教師と特権学生の政策の間のトークンレベルの監督を動的にルーティングするアドバンテージ・アウェアな二重蒸留パラダイムであるDOPDを提案する。
大きな言語モデル(LLM)と視覚言語モデル(VLM)の両方の実験は、DOPDがバニラPDを一貫して上回っていることを示している。
論文 参考訳(メタデータ) (2026-06-29T17:55:53Z) - EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation [5.310892696470208]
On-Policy Distillation (OPD)はLLMポストトレーニングパラダイムとして広く注目を集めている。
このアプローチの課題は、特権情報によって、意図よりもモデル行動を変えることができることだ。
EviDence GuidEd On-Policy Distillation (EDGE-OPD)を提案する。
論文 参考訳(メタデータ) (2026-05-22T10:55:15Z) - Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning [41.384652481442735]
我々は,一様教師模倣からエントロピー制御された指向性監視へと特権的な自己蒸留を再構成するtextbfDirection-Adaptive Self-Distillation (textbfDASD)を提案する。
6つの数学的推論ベンチマークで、DASDは強力なRLVRと自己蒸留ベースラインよりも優れたマクロAvg@16を達成する。
論文 参考訳(メタデータ) (2026-05-21T10:07:46Z) - Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation [49.117085054884676]
オンライン蒸留は、より強い教師からの強いフィードバックを使って、学生モデルを独自のロールアウトで訓練する。
我々は、この原則を軌跡固有のリリースルールで運用する。
強弱蒸留作業による実験結果から, この放出規則は標準全軌道PDよりも一貫して優れていたことが示唆された。
論文 参考訳(メタデータ) (2026-05-13T15:05:30Z) - The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes [10.319573084070578]
オンライン蒸留(OPD)とオンライン自己蒸留(OPSD)は,大規模言語モデルのための有望なポストトレーニング手法として出現している。
我々は、OPDとOPSDがいつ機能するか、いつ機能しないのか、なぜ機能しないのかについて、総合的な実証的研究を行った。
論文 参考訳(メタデータ) (2026-05-11T19:44:59Z) - TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents [55.27396165691312]
マルチターンエージェント設定におけるバニラOPDの鍵となる制限を,トラジェクトリレベルKL不安定(Trajectory-Level KL Instability)と呼ぶ。
学生に露出する軌道深度を制御し,カリキュラムのスケジュールを段階的に拡張するフレームワークであるTCODを提案する。
4組の生徒と教師のペアによる実験結果から,TCODはKLのエスカレーションを軽減し,トレーニングを通してKLの安定性を高め,バニラPDよりも最大18ポイントのエージェント性能を向上させることが示された。
論文 参考訳(メタデータ) (2026-04-27T03:38:27Z) - Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models [44.041109669153506]
On-Policy Self-Distillation (OPSD) は、教師と学生の両方がひとつのモデルで、異なるコンテキストを条件付けして機能するフレームワークである。
複数の数学的推論ベンチマークにおいて,本手法の有効性を示す。
論文 参考訳(メタデータ) (2026-01-26T17:56:50Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。