論文の概要: d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
- arxiv url: http://arxiv.org/abs/2609.35362v1
- Date: Mon, 28 Sep 2026 15:08:22 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-02 06:11:38.002546
- Title: d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
- Title(参考訳): d-OPD:ブロック拡散言語モデルのための将来性を考慮したオンライン蒸留
- Abstract要約: 我々は,AR教師の分布を改良し,生徒の目に見える状態に適合する将来的なオンライン蒸留法であるd-OPDを紹介した。
Qwen3の0.6Bから8Bまでのモデル全体で、d-OPDはOPDLMよりも最大4.0ドルポイント向上している。
- 参考スコア(独自算出の注目度): 12.972376712003287
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to $4.0$ points over OPDLM and reduces training time by $1.35$-$1.58\times$. The code is available at https://github.com/mit-han-lab/d-OPD.
- Abstract(参考訳): 大規模言語モデル(LLM)は通常、テキストの自動回帰(AR)を生成し、一度に1つのトークンを予測する。
ブロック拡散言語モデル(dLLM)は、ブロックを逐次生成する代わりに、各ブロック内で複数のトークンを並列に宣言し、生成を加速する有望な方法を提供する。
このようなモデルをゼロからトレーニングするのではなく、最近の研究は、強い事前訓練されたARモデルを蒸留によりブロックdLLMに適応させる。
オンライン蒸留(OPD)は、固定されたオフライン軌道のみに限らず、現在の政策によって生成される状態において学生を監督するため、LLM訓練に広く用いられている。
学生が実際に訪れている州をトレーニングすることで、トレーニングと世代間のミスマッチを減らし、学生が進化するにつれてより適切な監督を提供することができる。
最近の研究により、このアイデアはAR-to-block-diffusion変換に拡張されている。
しかし、この設定では、ブロック拡散学生と因果AR教師が、同じ訓練状態の異なる情報について、基本的なミスマッチを導入している。
学生は、将来の可視的文脈を含む部分的認知ブロック全体から予測し、標準的なAR教師ターゲットは、因果接頭辞からのみ定義される。
その結果、蒸留に使用する教師分布は、学生が利用できる情報と完全に一致していない。
そこで我々は,各ブロックに目に見える将来の情報を組み込むことで,AR教師の配当を改良し,学生が使用する情報とよりよく一致するような指導を行う,将来認識型オンライン蒸留法であるd-OPDを導入する。
Qwen3 モデルは 0.6B から 8B までで、d-OPD は OPDLM で最大 4.0 ポイントまで改善し、トレーニング時間を $1.35$-1.58\times$ に短縮する。
コードはhttps://github.com/mit-han-lab/d-OPD.comで公開されている。
関連論文リスト
- ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation [12.983899693598396]
拡散言語モデル(DLM)はフレキシブルトークン順序と並列生成を提供する。
直接蒸留は基本的なミスマッチに直面し、自己回帰的な教師は左の接頭辞から予測し、一方DLMは両側のトークンに条件を付ける。
本稿では,教師の指導から学生のロールアウトを分離することで,このミスマッチを解決する蒸留フレームワークであるForkLeftを紹介する。
論文 参考訳(メタデータ) (2026-09-26T10:29:45Z) - Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation [97.02922768457239]
Context-Matched Distillation (CMD) は、教師の監督と、各ターゲットが生成される際に利用可能な情報とを整合させる因果的MDDフレームワークである。
Prefix Corruptionは、このターゲットコンテキストアライメントを維持しながら、トレーニングの初期段階で生成された信頼性の低いプレフィックスを摂動することで、トレーニングを安定化させる。
ショートビデオとロングビデオの両方のベンチマークで、自己回帰手法における最先端の集約性能を実証した。
論文 参考訳(メタデータ) (2026-08-13T15:51:31Z) - Weak-to-Strong On-Policy Distillation [21.651421588571296]
Weak-to-Strong On-Policy Distillation (W2S-OPD)を導入した。
4つの数学と3つのコードベンチマークで、W2S-OPDはOPDを上回り、学生がドメイン教師を超越し、監督源が弱くなっても改善を続けることができる。
論文 参考訳(メタデータ) (2026-07-28T20:31:48Z) - H$^2$SD: Hybrid Hindsight Self-Distillation [61.71131241473028]
検証可能な報酬付き強化学習(RLVR)は、言語モデル推論のための信頼性の高い結果監視を提供する。
既存の自己蒸留法は特権教師を追加するが、通常は一定の役割を割り当てる。
本稿では,教師のコンテキストに適応し,トラジェクタの正当性を更新するHybrid Hindsight Self-Distillation(MathrmH2mathrmSD$)を紹介する。
論文 参考訳(メタデータ) (2026-07-21T10:47:27Z) - Trace-Based On-Policy Distillation for Masked Diffusion Language Models [14.931427442072502]
dLLMの監視された微調整 (SFT) には、密度が高いが、しばしば非政治的なマスク状態が必要である。
強化学習(RL)はスパース報酬や価値モデリングに依存している。
本稿では, 教師が指導する目標dLLMに推論能力を伝達するフレームワークであるtextbftrace-based on-policy distillation (TOPD)を提案する。
論文 参考訳(メタデータ) (2026-07-18T16:25:17Z) - DemoPSD: Disagreement-Modulated Policy Self-Distillation [49.30809363804017]
大規模言語モデルを訓練するための実践的な方法として、オンデマンド自己蒸留が登場している。
DemoPSDは、教師と生徒の分布の重み付けされた幾何学的組み合わせに向けて学生を操縦する。
論文 参考訳(メタデータ) (2026-07-02T17:58:29Z) - Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion [84.20493238687187]
単純な模倣を超えて、より深い概念的理解を具現化する新しい枠組みを導入する。
underlinetextitFirst, to address pattern memorization, Explanatory Inversion (EI) generated target explanatory probes'
underlinetextitSecondは、一般化を改善するために、Explainatory GRPO (texttEXGRPO) は、新しいダイアログ構造ユーティリティーボーナスを用いた強化学習アルゴリズムを使用する。
論文 参考訳(メタデータ) (2026-02-26T23:01:46Z) - From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs [58.640039233470766]
原理的AR-to-block-diffusion適応は,DLMをスクラッチからトレーニングする上で,有効かつ効率的な代替手段であることを示す。
NBDiff-7B(BaseとInstruct)は、長文のモデリングと推論機能を継承し、最先端のパフォーマンスを実現する。
論文 参考訳(メタデータ) (2025-12-07T10:28:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。