論文の概要: Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
- arxiv url: http://arxiv.org/abs/2610.02700v1
- Date: Fri, 02 Oct 2026 02:34:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.167128
- Title: Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
- Title(参考訳): 進化するエラーから学ぶ:オンライン蒸留における適応的反復修復
- Abstract要約: オンラインの自己蒸留は、学生自身の方針からサンプリングされた軌跡に対する密集したトークンレベルのフィードバックを提供する。
本稿では, オンライン蒸留における適応的反復補修フレームワークであるAIR-OPDを紹介する。
DAPO-Math-17Kデータセット上でAIR-OPDをトレーニングし, AIME24, AIME25, HMMT25, MMLU-Pro, GPQAのアウト・オブ・ディストリビューションテストと並行して評価する。
- 参考スコア(独自算出の注目度): 34.25144986131144
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
- Abstract(参考訳): On-policy Self-distillation (OPSD) は、学生自身の方針から得られたトラジェクトリに対する高密度トークンレベルのフィードバックを提供する。
このフィードバックは、学生が利用できない完全な参照ソリューションで条件付けられた教師から得られます。
参照ソリューションは、ターゲットを特定するが、生徒の現在のエラーからそれへの移動方法を規定せず、ソリューション条件のショートカットリスクを生成する。
本稿では, オンライン蒸留の適応的反復的修復フレームワークであるAIR-OPDを紹介する。
応答が失敗すると、誘導発生器は、現在のエラーに対する修理誘導を合成する。
学生は、このガイダンスで政治上の再試行をサンプリングします。
再試行が不正確であれば、発電機は新たに観測されたエラーに対する新しい修理ガイダンスを生成する。
各ラウンドにおいて、固定教師は、特権的文脈として指導を受け、最新の失敗応答のエラー整合領域で生徒を監督する。
アウトカム・アウェアのステージ重み付けは、早期の修復段階と、即時再試行が検証に合格するクレジットステージを優先する。
DAPO-Math-17Kデータセット上でAIR-OPDをトレーニングし, AIME24, AIME25, HMMT25, MMLU-Pro, GPQAのアウト・オブ・ディストリビューションテストと並行して評価する。
本研究では、現在の学生政策からの自己指導と、より大きなモデルからの外部指導の2つのガイダンス源について検討する。
Qwen3-4BとQwen3-8Bでは、AIR-OPDは最高の数学的推論平均に達し、最強のベースラインを最大3.6ポイント改善した。
関連論文リスト
- WAM-OPD: Sharpening World Action Models via On-Policy Distillation [84.1328443874472]
事前訓練された世界行動モデル(WAM)は、多様なロボット操作タスクにまたがる汎用機能を提供する。
WAMのオンライン蒸留(OPD)とWAM-OPDの導入について検討する。
We show WAM-OPD improves target-task performance while maintaining near-itial performance on tasks excludeed by adapt。
論文 参考訳(メタデータ) (2026-09-28T03:56:23Z) - When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation [21.86522736658343]
本稿では, 教師指導の誤りを軽減するために, Reward-Aligned On-Policy Distillation (RA-OPD)を提案する。
RA-OPDは、さらなる計算コストを必要とせずに、より信頼性の高い軌道を選択し、生徒モデルの性能を向上させる。
我々は,Qwen3ファミリーとDeepSeek-R1ファミリーのモデルを用いて,数学とコードベンチマークのRA-OPDを評価する。
論文 参考訳(メタデータ) (2026-08-28T06:05:50Z) - H$^2$SD: Hybrid Hindsight Self-Distillation [61.71131241473028]
検証可能な報酬付き強化学習(RLVR)は、言語モデル推論のための信頼性の高い結果監視を提供する。
既存の自己蒸留法は特権教師を追加するが、通常は一定の役割を割り当てる。
本稿では,教師のコンテキストに適応し,トラジェクタの正当性を更新するHybrid Hindsight Self-Distillation(MathrmH2mathrmSD$)を紹介する。
論文 参考訳(メタデータ) (2026-07-21T10:47:27Z) - A Formula-Driven Survey and Research Agenda for On-Policy Distillation [4.397842507533513]
本調査では,OPDを単一損失ファミリーではなく,フィードバックから更新までの問題として検討した。
我々は, 直接分布損失と政策段階の対数比更新という2つの経路から公式駆動型分類法を開発した。
論文 参考訳(メタデータ) (2026-06-22T03:09:21Z) - When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks [2.743683637024251]
検証者駆動型自己DPOは、自己改善型視覚言語モデルのための一般的なレシピである。
検証器の品質がタスク固有のため,この仮定は失敗する可能性がある。
本稿では,プログレッシブゲートリプレイとその方向ミスマッチ故障モードに対する分散定理を用いて,コンパクトなメカニスティックな説明を行う。
論文 参考訳(メタデータ) (2026-06-12T16:55:30Z) - INFUSER: Influence-Guided Self-Evolution Improves Reasoning [54.101135873140066]
2つの共進化的役割を持つ反復的協調学習フレームワークを導入する。
解答器は、生成元が提供する回答に対して標準正当性報酬で訓練される。
8B INF共進化ジェネレータは、数学とコーディングにおいて凍った32B思考ジェネレータより優れている。
論文 参考訳(メタデータ) (2026-06-08T05:40:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。