Fugu-MT 論文翻訳(概要): Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

論文の概要: Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arxiv url: http://arxiv.org/abs/2605.03677v1
Date: Tue, 05 May 2026 12:15:21 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-06 19:35:43.925552
Title: Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Title（参考訳）: Uni-OPD:デュアルパースペクティブレシピによるオンポリシィ蒸留の統合
Authors: Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan,
Abstract要約: LLM(Large Language Models)とMLLM(Multimodal Large Language Models)をまたいで一般化する統一OPDフレームワークであるUni-OPDを提案する。具体的には、学生の立場から、学習中の情報発信状態の探索を促進するために、2つのデータバランス戦略を採用する。我々は,正しい軌道と間違った軌道の順序の整合性を取り戻すために,結果誘導マージンキャリブレーション機構を開発した。
参考スコア（独自算出の注目度）: 53.40076304466524
License: http://creativecommons.org/licenses/by/4.0/
Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the conditions under which OPD yields reliable improvement remain poorly understood. In this work, we identify two fundamental bottlenecks that limit effective OPD: insufficient exploration of informative states and unreliable teacher supervision for student rollouts. Building on this insight, we propose Uni-OPD, a unified OPD framework that generalizes across Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), centered on a dual-perspective optimization strategy. Specifically, from the student's perspective, we adopt two data balancing strategies to promote exploration of informative student-generated states during training. From the teacher's perspective, we show that reliable supervision hinges on whether aggregated token-level guidance remains order-consistent with the outcome reward. To this end, we develop an outcome-guided margin calibration mechanism to restore order consistency between correct and incorrect trajectories. We conduct extensive experiments on 5 domains and 16 benchmarks covering diverse settings, including single-teacher and multi-teacher distillation across LLMs and MLLMs, strong-to-weak distillation, and cross-modal distillation. Our results verify the effectiveness and versatility of Uni-OPD and provide practical insights into reliable OPD.
Abstract（参考訳）: オンライン蒸留(OPD)は、最近、専門的専門家モデルの能力を単一学生モデルに統合するための効果的なポストトレーニングパラダイムとして登場した。実証的な成功にもかかわらず、OPDが信頼できる改善をもたらす条件はよく分かっていない。本研究では,効果的なOPDを制限する2つの基本的なボトルネック,すなわち,情報的状態の探索が不十分なことと,学生のロールアウトに対する教師の信頼性の低いことを明らかにする。この知見に基づいて,LLM(Large Language Models)とMLLM(Multimodal Large Language Models)をまたいで一般化する統一OPDフレームワークUni-OPDを提案する。具体的には、学生の立場から、学習中の情報発信状態の探索を促進するために、2つのデータバランス戦略を採用する。教師の立場から,集計トークンレベルの指導が結果報酬と整合性を維持しているかどうかを,信頼性の高い監督が判断することを示す。そこで本研究では,正しい軌道と不正確な軌道の整合性を復元する結果誘導マージンキャリブレーション機構を開発した。筆者らは, LLM, MLLM, 強弱蒸留, クロスモーダル蒸留など, 5 つのドメインと16 のベンチマーク実験を行った。本研究は,Uni-OPDの有効性と汎用性を検証し,信頼性の高いOPDに関する実践的な洞察を提供する。

論文の概要: Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

関連論文リスト