論文の概要: MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
- arxiv url: http://arxiv.org/abs/2607.17751v1
- Date: Mon, 20 Jul 2026 09:44:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.57723
- Title: MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
- Title(参考訳): MagicSelector:Counterfactal DecompositionとProgressive Re rankによるエージェントツール選択のための共同最適化
- Authors: HONOR Agentic Search Team, Zhengzong Chen, Lei Tang, Lijun Liu, Chuandi Jiang, Fan Yang, Keyun Chu, Chu Zhao, Shihao Liu, Minghang Li, Bo Liang, Can Wen, Hailong Wu, Jingnan Ju, Mian Liu, Nengbin Zhang, Peiqiang Wang, Penghe Nie, Qinhui Gu, Sijia Lv, Siqi Chen, Wei Zhang, Yang Xu, Yuhao Qian, Yuxiang Zhang, Zeng Cheng, Zhen Wang, Zuan Chen, Yuanyuan Zhao, Fei Huang,
- Abstract要約: MagicSelectorは、あいまいなユーザー命令を実行可能なアトミックサブタスクに変換することができるフレームワークである。
ドメイン外(OOD)シナリオにおける高精度ツール検索をガイドする。
- 参考スコア(独自算出の注目度): 31.430385887150972
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios.We empower MagicSelector with these capabilities through three key contributions: (1) a preferenceguided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.
- Abstract(参考訳): エージェントにおけるツール検索の基本的な課題に対処するために,非現実的タスク分解,プログレッシブリグレード,動的トップKを統合した共同最適化フレームワークMagicSelectorを提案する。
MagicSelectorは、不明瞭なユーザ命令を実行可能なアトミックサブタスクに翻訳し、高精度ツール検索を誘導し、ドメイン外(OOD)シナリオにおける冗長なノイズや厳しいコンテキストの中断を効果的に軽減する、特殊なフレームワークである。我々はMagicSelectorに3つの重要な貢献を通じて、これらの機能を強化する。(1)逆ファクト的タスク分解機構を利用して、検索ランキングにおける分解の限界因果的利益を定量化し、論理的コヒーレンスに効果的にきめ細かな構造的監督を効果的に付与するプログレッシブツールリグレート手法、(2)類似ツール間の細かな識別を最適化するポイントワイドとリストワイドの関連性を最適化するプログレッシブツールリグレート手法、(3)2つのセマンティック・セマンティック・セマンティック・セマンティック・セマンティック・セマンティック・セマンティック・セマンティック・セマンティクスのスコアを動的に監視する機能。
MTDToolは,プロセスレベルのアノテーションを用いたモバイルマルチターンインタラクションに適した,最初のタスク分解ベンチマークです。
その結果、MagicSelectorはツール検索精度、OOD一般化能力、トークン全体の効率において最先端の手法よりも優れており、提案フレームワークの有効性が示された。
関連論文リスト
- PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval [32.19528642807244]
タスク分解は曖昧な命令を実行可能なアトミックサブタスクに変換することを目的としている。
タスク分解に対する報酬は、強化学習に基づく手法で容易に報酬のハッキングを誘発することができる。
提案手法はPCTD, PCTD, Preference-guided Counterfactual Task Decomposition である。
論文 参考訳(メタデータ) (2026-07-17T07:15:41Z) - From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents [90.3279207163486]
EvoSOPは、エージェントが実行軌跡からSOPを抽出し、反復的にツールセットを最適化することを可能にするフレームワークである。
EvoSOPはタスク成功率を大幅に向上し,インタラクションラウンドの回数を大幅に削減することを示す。
論文 参考訳(メタデータ) (2026-07-08T12:09:49Z) - Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models [27.250148827297604]
HDPOは、ツールの効率を競合するスカラー目標から厳格な条件に書き換えるフレームワークです。
私たちのモデルであるMetisは、推論精度を同時に高めながら、ツールの呼び出しを桁違いに削減します。
論文 参考訳(メタデータ) (2026-04-09T17:59:57Z) - Adaptive Robust Estimator for Multi-Agent Reinforcement Learning [27.595086716369483]
協調推論のための頑健な多エージェント強化学習フレームワークを提案する。
Dual-Agent Answer-Critique-Rewrite (DACR)とAdaptive Robust Estimator (ARE)の2つのコンポーネントで構成されている。
論文 参考訳(メタデータ) (2026-03-23T04:51:15Z) - Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents [54.18201810286764]
LLM(Large Language Models)に基づくツール利用エージェントは、数学的推論やマルチホップ質問応答といったタスクに優れる。
長い道のりでは、エージェントはしばしば過度で低品質なツールコールをトリガーし、レイテンシを増大させ、推論性能を低下させる。
本稿では,エントロピー低減を監視信号として使用し,ツール使用行動の最適化ニーズに対処する2つの報奨戦略を設計する。
論文 参考訳(メタデータ) (2026-02-02T12:52:14Z) - PerfGuard: A Performance-Aware Agent for Visual Content Generation [53.591105729011595]
PerfGuardは、ビジュアルコンテンツ生成のためのパフォーマンス対応のエージェントフレームワークである。
ツールのパフォーマンス境界をタスク計画とスケジューリングに統合する。
ツール選択の正確性、実行の信頼性、ユーザの意図との整合性にメリットがあります。
論文 参考訳(メタデータ) (2026-01-30T05:12:19Z) - On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows [71.92083784393418]
エージェントAI(自律的な計画と行動を行うシステム)は広く普及しているが、複雑なタスクにおけるタスクの成功率は低いままである。
推論時のアライメントは、サンプリング、評価、フィードバックの3つのコンポーネントに依存します。
本稿では,様々な形態の批判から抽出されたフィードバックを繰り返し挿入するIterative Agent Decoding(IAD)を紹介する。
論文 参考訳(メタデータ) (2025-04-02T17:40:47Z) - AdaRank: Adaptive Rank Pruning for Enhanced Model Merging [23.649762835129167]
モデルマージは、独立して微調整されたモデルを統合されたフレームワークに統合するための有望なアプローチとして現れている。
AdaRankは、タスクベクトルの最も有用な特異な方向を適応的に選択し、複数のモデルをマージする新しいモデルマージフレームワークである。
AdaRankは、さまざまなバックボーンとタスク数で一貫して最先端のパフォーマンスを実現し、微調整されたモデル間のパフォーマンスギャップを1%近く削減している。
論文 参考訳(メタデータ) (2025-03-28T06:49:06Z) - Meta-Wrapper: Differentiable Wrapping Operator for User Interest
Selection in CTR Prediction [97.99938802797377]
クリックスルー率(CTR)予測は、ユーザーが商品をクリックする確率を予測することを目的としており、リコメンデーションシステムにおいてますます重要になっている。
近年,ユーザの行動からユーザの興味を自動的に抽出する深層学習モデルが大きな成功を収めている。
そこで我々は,メタラッパー(Meta-Wrapper)と呼ばれるラッパー手法の枠組みに基づく新しい手法を提案する。
論文 参考訳(メタデータ) (2022-06-28T03:28:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。