論文の概要: From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- arxiv url: http://arxiv.org/abs/2607.24280v1
- Date: Mon, 27 Jul 2026 11:27:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.402076
- Title: From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- Title(参考訳): プロプライエタリからオープンソースへ:エージェントサーチにおけるマルチエージェントプロトコル蒸留による配電ギャップのブリッジ
- Authors: Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou,
- Abstract要約: エージェントサーチにより、多段階推論と検索をインターリーブすることで、大規模言語モデルによる知識集約的なタスクの解決が可能になる。
知識蒸留は指導を与えることができ、強力な推論能力を持つ先進的なプロプライエタリなモデルは有望な教師である。
共同蒸留と結果に基づく強化学習フレームワークであるMAPD(Multi-Agent Protocol Distillation)を提案する。
- 参考スコア(独自算出の注目度): 12.160113054571356
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded by hidden logits and mismatched tokenizers, whereas raw natural language trajectory imitation transfers superficial stylistic artifacts rather than core reasoning competence. To address the heterogeneous distillation problem and bridge the distribution gap, we propose Multi-Agent Protocol Distillation (MAPD), a joint distillation and RL framework uses a structured, style-normalized protocol as an intermediate representation. An offline multi-agent system (MAS) decomposes each query, retrieves supporting evidence, repairs failed searches, and converts the resulting exploration trace into a JSON protocol containing the task type, reasoning plan, and extractive grounding facts. During training, the protocol is provided only to a privileged branch of the student policy, whose token distributions furnish a dense distillation signal alongside the sparse RL objective. Extensive evaluations across seven QA benchmarks demonstrate that MAPD consistently outperforms competitive distillation and RL, achieving average success rates of 39.4\% on Qwen3-1.7B and 44.4\% on Qwen3-4B. Crucially, the framework generalizes robustly across diverse proprietary teachers while effectively mitigating the student policy from style drift and verbosity degeneration.
- Abstract(参考訳): エージェント検索により,多段階推論と検索をインターリーブすることで,大規模言語モデルによる知識集約的な課題の解決が可能となる。
知識蒸留はより密集したガイダンスを提供することができ、強力な推論能力を持つ先進的なプロプライエタリなモデルは有望な教師である。
プロプライエタリなモデルからの蒸留は、このスーパーバイザリーシグナルを密度化することができるが、従来のロジットマッチングは、隠されたロジットや不正なトークン化業者によって妨げられ、一方、生の自然言語の軌道模倣は、コア推論能力ではなく表面的なスタイル的アーティファクトを伝達する。
ヘテロジニアス蒸留問題に対処し, 分配ギャップを埋めるために, 中間表現として構造化されたスタイル正規化プロトコルを用いた多エージェントプロトコル蒸留(MAPD)を提案する。
オフラインマルチエージェントシステム(MAS)は、各クエリを分解し、エビデンスを検索し、失敗した検索を修復し、結果の探索トレースをタスクタイプ、推論計画、抽出根拠事実を含むJSONプロトコルに変換する。
訓練中、このプロトコルは学生ポリシーの特権分岐にのみ提供され、そのトークン分布はスパースRL目的と並んで密度の高い蒸留信号を付与する。
7つのQAベンチマークにおいて、MAPDは競争蒸留とRLを一貫して上回り、Qwen3-1.7Bの平均成功率は39.4\%、Qwen3-4Bの平均成功率は44.4\%に達することを示した。
重要な点として、このフレームワークは、様々なプロプライエタリな教師に対して堅牢に一般化し、学生政策を、スタイルドリフトと冗長性変性から効果的に緩和する。
関連論文リスト
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models [55.01951088768769]
DiffusionOPDはオンライン政策蒸留(OPD)に基づく拡散モデルのための新しいマルチタスクトレーニングパラダイムである
本研究では,DiffusionOPDがトレーニング効率と最終性能において,マルチリワードRLとカスケードRLのベースラインを一貫して上回っていることを示す。
論文 参考訳(メタデータ) (2026-05-14T16:49:09Z) - AutoResearch-RL: Perpetual Self-Evaluating Reinforcement Learning Agents for Autonomous Neural Architecture Discovery [5.110708177092157]
本稿では、強化学習エージェントが人間の監督なしにオープンエンドニューラルネットワーク研究を行うためのフレームワークであるAutoResearch-RLを提案する。
我々はこれをマルコフ決定過程として定式化し、軽微な仮定の下で収束保証を導出し、1つのGPUナノチャット事前学習ベンチマークで経験的に実証する。
論文 参考訳(メタデータ) (2026-03-07T17:49:44Z) - Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration [49.9937230730202]
本稿では,新たなアクター・リファイナ・コラボレーション・フレームワークであるSearch-R2を提案する。
提案手法は,生成過程をアクターに分解し,最初の推論軌道を生成する。
本稿では,検索-R2がモデルスケール全体にわたって強力なRAGとRLベースのベースラインを一貫して上回ることを示す。
論文 参考訳(メタデータ) (2026-02-03T15:32:09Z) - Coupled Variational Reinforcement Learning for Language Model General Reasoning [83.82392089177841]
変分推論と強化学習を橋渡しするために,textitbCoupled bVari bReinforcement bLearning (CoVRL)を提案する。
CoVRLはベースモデルよりも12.4%向上し、最先端の検証不要なRLベースラインよりも2.3%向上した。
論文 参考訳(メタデータ) (2025-12-14T07:03:51Z) - Multimodal Reinforcement Learning with Agentic Verifier for AI Agents [131.46008226323423]
Argosは、エージェントタスクの推論モデルをトレーニングするための、原則化されたマルチモーダル報酬エージェントである。
エージェント検証をSFTデータとRLトレーニングの両方で活用することにより、我々のモデルは最先端の結果を得ることができる。
論文 参考訳(メタデータ) (2025-12-03T04:42:47Z) - MALT: Improving Reasoning with Multi-Agent LLM Training [67.76186488361685]
MALT(Multi-Agent LLM Training)は、推論プロセスを生成、検証、改善ステップに分割する、新しいポストトレーニング戦略である。
MATH、GSM8K、CSQAでは、MALTは、それぞれ15.66%、7.42%、9.40%の相対的な改善で同じベースラインLLMを上回っている。
論文 参考訳(メタデータ) (2024-12-02T19:30:36Z) - Deep Multi-Agent Reinforcement Learning for Decentralized Active
Hypothesis Testing [11.639503711252663]
我々は,深層多エージェント強化学習の枠組みに根ざした新しいアルゴリズムを導入することで,マルチエージェント能動仮説テスト(AHT)問題に取り組む。
エージェントが協調戦略を学習し、性能を向上させる能力を効果的に示す実験結果を包括的に提示する。
論文 参考訳(メタデータ) (2023-09-14T01:18:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。