論文の概要: Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
- arxiv url: http://arxiv.org/abs/2608.06171v1
- Date: Thu, 06 Aug 2026 15:37:04 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-07 15:25:20.944943
- Title: Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
- Title(参考訳): Webエージェントの表現ルーティングは、最も価値の高い場所で最も学習しやすい
- Authors: Jiaming Wei, Zekun Wu, Adriano Koshiyama, Maria Perez-Ortiz,
- Abstract要約: VisualWebArenaとWebArenaの8つのサイトモデルの組み合わせ(セル)の6つの観察モードを測定した。
次に、5つのルーティングポリシー(モードの選択、強いモードにいつ費やすかの決定、タスクテキストを読み取るゼロコストルール、信頼性カスケード、プールされたコスト層)をテストする。
たった1つのウェル・チョーゼン・モードを固定するだけじゃなく、たった1つの例外は、我々の最短の細胞に対する脆弱な結果です。
- 参考スコア(独自算出の注目度): 0.8491218071542773
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others miss, they fail in structurally different ways, and the best choice reverses between task sets. The obvious prize, an oracle that picks a winning mode for every task, looks large but is inflated by run-to-run noise: rerunning the same mode on the same tasks changes 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. What survives is a cost bound: sending only the tasks no mode solves to the cheapest mode cuts cost by 9.5-30.6% in 8 of 8 cells at unchanged success. We then test five routing policies (picking the mode, deciding when to spend on the strong mode, a zero-cost rule read off the task text, a confidence cascade, and pooled cost tiers), and none robustly beats simply fixing one well-chosen mode; the one exception is a fragile result in our sparsest cell. The central obstruction is that routing supervision is produced at the agent's success rate: the weaker the agent, the fewer labels a router gets, exactly where routing would be most valuable. This limit belongs to today's agents rather than to routing itself. Label supply and routing opportunity rise together (correlation 0.95 across cells), so a stronger agent can overturn the result, and we report the rerun noise bands and the full measurement protocol.
- Abstract(参考訳): Webエージェントは、テキスト、ピクセル、または両方を通じてブラウザを観察し、選択は通常、すべてのタスクに対して1度だけ修正される。
VisualWebArenaとWebArenaの8つのサイトモデルの組み合わせ(セル)にまたがる6つの観察モードを測定し、タスク毎に何を買うかを尋ねる。
各モードは相補的であり、それぞれが他のタスクを見逃し、構造的に異なる方法で失敗し、最適な選択はタスクセット間で逆転する。
同じタスクで同じモードを実行すると、結果の12~14%が変わるので、既に手元にあるモードの2回目の実行は、新しいモードを追加するのと同じくらい多くなります。
モードなしのタスクのみを最も安価なモードに送信することで8セル中8セルで9.5-30.6%のコスト削減が可能となる。
次に、5つのルーティングポリシー(モードの選択、強いモードの時間の決定、タスクテキストを読み取るゼロコストルール、信頼性カスケード、プールされたコスト階層)をテストする。
中央の障害は、ルーティングの監督がエージェントの成功率で生成されることである:エージェントが弱ければなるほど、ルータが取得するラベルが少なくなる。
この制限は、ルーティング自体ではなく、今日のエージェントに属します。
ラベルの供給とルーティングの機会が一緒になり(細胞間の相関0.95)、より強いエージェントが結果を覆い、リランノイズバンドと全測定プロトコルを報告する。
関連論文リスト
- Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First [0.15293427903448023]
リポジトリをスカウトした後にルートするSuperScoutを紹介します。
7B検索エンジンであるSuperScout-7Bは、まずリポジトリを探索し、構造化されたハンドオフを生成する。
検索者の隠された状態とタスクテキストは、履歴ベースのルータを送信し、4つのフロンティアフィクスラーのうちの1つにタスクをディスパッチする。
論文 参考訳(メタデータ) (2026-08-05T13:11:31Z) - TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI [8.291830023781403]
本稿では,タスクレベルのルーティングフレームワークであるTRACEについて述べる。
タスクの遅延フィードバックを活用することで、TRACEはワークロードに適応するルーティングポリシを学ぶ。
精度にマッチしたトレードオフを継続的に改善し、非レイテンシフロンティアポイントを達成します。
論文 参考訳(メタデータ) (2026-07-24T16:29:06Z) - Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction [57.138342889101345]
Tencent WorkBuddy Benchは、コーディングエージェントのためのマルチドメイン評価スイートである。
本報告では, 設計手法, スコアリングプロトコル, クロスモデルリーダーボードについて述べる。
論文 参考訳(メタデータ) (2026-07-23T04:34:06Z) - Agentic Routing: The Harness-Native Data Flywheel [35.444792826526]
DRACOやPinchBenchなどのエージェントベンチマークにおけるシングルトンおよびマルチモデルルーティングに関する研究
エージェントルーティングは単なるコストコントロールではなく、エージェントネイティブなトレーニングのためのデータエンジンである、と氏は主張する。
論文 参考訳(メタデータ) (2026-07-13T11:05:55Z) - Agent-as-a-Router: Agentic Model Routing for Coding Tasks [54.00153674547733]
Agent-as-a-router (AC)はC-A-Fループとしてルーティングを形式化する(Context->Action->FeedbackContext)
ACは、分配タスクに対する最小の累積後悔を達成し、配布外エージェントプログラミングタスクに一般化する。
論文 参考訳(メタデータ) (2026-06-22T06:37:31Z) - When Routing Collapses: On the Degenerate Convergence of LLM Routers [46.01380774114097]
ユーザのコスト予算が増加するにつれて、ルータは体系的に最も有能で最も高価なモデルにデフォルトとなる。
モデルランキングを直接学習する決定対応ルータであるEquiを提案する。
RouterBenchでは、最強の先行ルータと比較して、GPT-4レベルのパフォーマンスでコストを約17%削減する。
論文 参考訳(メタデータ) (2026-02-03T12:51:55Z) - BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks [51.803138848305814]
我々はBrowserArenaを紹介した。BrowserArenaは、ユーザから送信されたタスクを収集するオープンソースのエージェント評価プラットフォームである。
Captcha解決、ポップアップバナー削除、URLへのダイレクトナビゲーションの3つの一貫した障害モードを特定します。
本研究は,Webエージェントの多様性と脆性の両方を明らかにする。
論文 参考訳(メタデータ) (2025-10-02T15:22:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。