論文の概要: Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
- arxiv url: http://arxiv.org/abs/2608.01953v2
- Date: Wed, 05 Aug 2026 08:08:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.135172
- Title: Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
- Title(参考訳): 飲酒前を振り返る: エージェントオンポリシィ蒸留のための教師指導の今後の軌道検証
- Authors: Chishui Chen, Yaoyou Fan, Te Sun, Yi Yang, Chenghao Sun, Delin Mao, Hongbo Qiao, Zuowei Zhang, Junxi Wang, Chenxing Sun, Yangen Hu, Lu Pan, Xuyang Liu, Linfeng Zhang,
- Abstract要約: オンライン蒸留(On-policy distillation、OPD)は、学生が訪れた州を教師が監督する制度である。
本稿では,教師の短い橋渡しを行うFutureBridge-OPD(FTB)を提案する。
Qwen3-32Bの主教師がQwen3-1.7Bの学生設定で、FTBはバニラOPDとTCODを平均16.6点、7.6点で上回っている。
- 参考スコア(独自算出の注目度): 18.893534946567218
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.
- Abstract(参考訳): オンライン蒸留(On-policy distillation、OPD)は、学生が訪れた州を教師が監督し、訓練と推論の間の分配ギャップを減らす。
しかし、多ターンエージェントタスクでは、学生の偏差は時間とともに蓄積し、教師指導が有効である状態から徐々に軌跡を移動させる。
定量的分析により,高認知状態が教師指導に有望な機会を提供することが示されたが,そのような指導が有益かどうかを判断するには,その後の学生軌跡に対する影響を検討する必要がある。
本研究では,教師に対して短い橋梁を高い不一致で実行し,その結果の学生継続を利用して,橋が教師に対して正の蒸留信号の密度を増大させるかどうかを評価する。
ALFWorld、WebShop、ScienceWorldでは、Qwen3-32Bの教師がQwen3-1.7Bの学生設定で、FTBは、それぞれ16.6ポイントと7.6ポイントでバニラOPDとTCODを上回り、学生の規模と教師の設定で有効である。
私たちのコードはhttps://github.com/ChenChiShui/FutureBridge-OPD.comで公開されています。
関連論文リスト
- On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents [50.78688247330029]
オン・ポリシィ蒸留(On-Policy Distillation、OPD)は、小規模の学生にその能力を伝達するための自然なレシピである。
本稿では,教師と生徒が生成するターンを各ロールアウトに混在させるシンプルなアルゴリズムであるGuided-OPD(Guided-OPD)を提案する。
Guided-OPDは、平均してバニラOPDよりも21.1%、成功率25.5%向上している。
論文 参考訳(メタデータ) (2026-06-14T16:41:45Z) - On-Policy Distillation with Best-of-N Teacher Rollout Selection [54.91780727674628]
本報告では, オンライン蒸留のためのベスト・オブ・Nロールアウト教員選抜フレームワークBRTSを提案する。
BRTSは、教師軌道から構築された教師コンテキスト管理ブランチで、標準の学生コンテキストOPDを強化する。
BRTSは、挑戦的な推論ベンチマークにおいて、標準的なPDよりも改善されており、より難しいデータセットに対して最大の利益がある。
論文 参考訳(メタデータ) (2026-05-10T19:49:00Z) - TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents [55.27396165691312]
マルチターンエージェント設定におけるバニラOPDの鍵となる制限を,トラジェクトリレベルKL不安定(Trajectory-Level KL Instability)と呼ぶ。
学生に露出する軌道深度を制御し,カリキュラムのスケジュールを段階的に拡張するフレームワークであるTCODを提案する。
4組の生徒と教師のペアによる実験結果から,TCODはKLのエスカレーションを軽減し,トレーニングを通してKLの安定性を高め,バニラPDよりも最大18ポイントのエージェント性能を向上させることが示された。
論文 参考訳(メタデータ) (2026-04-27T03:38:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。