論文の概要: AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
- arxiv url: http://arxiv.org/abs/2606.13608v2
- Date: Sun, 14 Jun 2026 16:45:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-16 13:45:31.214862
- Title: AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
- Title(参考訳): エージェントBeats:オープンネス、標準化、再現性のためのエージェントアセスメント
- Abstract要約: エージェントシステムはドメイン間で急速に進歩しているが、その評価は断片化されている。
根本的問題は、オープンでエージェントに依存しないアセスメントインタフェースがないことである。
我々は、審査員が評価を行い、すべての参加者が標準化されたプロトコルを介して対話するエージェントエージェントアセスメント(AAA)を提唱する。
- 参考スコア(独自算出の注目度): 104.46861849039357
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.
- Abstract(参考訳): エージェントシステムはドメイン間で急速に進歩しているが、その評価は断片化されている。
ほとんどのベンチマークは、重い統合、テストプロダクションミスマッチの作成、さまざまなエージェント設計における公正な比較の制限を必要とする、固定されたLLM中心のハーネスに依存している。
根本的問題は、オープンでエージェントに依存しないアセスメントインタフェースがないことである。
我々は,エージェントエージェント評価(AA)を提唱し,審査員とすべての参加者が標準化されたプロトコルを通じて対話し,タスク管理のためのA2AとツールアクセスのためのMPPを提案する。
従来のベンチマークでは、ベンチマーク用とエージェント用という2つの別々のインターフェースが定義されているが、AAAでは1つしか必要とされていない。
我々はさらに,AAA の具体的実現として AgentBeats を導入し,オープン性,プライバシ,再現性に関する現実的な制約に適合する5つの実用的な運用モードを特定した。
大規模に評価するために、12のカテゴリで298人の審査員と467人の被験者エージェントを対象とする5ヶ月のオープンコンペティションを行い、AAAが不均一なベンチマークの範囲に適用されることを示した。
コミュニティスケールのフィールドスタディと制御されたコーディングケーススタディを組み合わせることで、AAAが大規模な異種シナリオに対してカバレッジ、実用性、忠実性を提供することを確認した。
AAAとAgentBeatsは、オープンで、標準化され、再現可能なエージェントアセスメントへの明確な道を提供する。
関連論文リスト
- PACE: A Proxy for Agentic Capability Evaluation [60.3743414796937]
PACEは、既存の非エージェント評価からインスタンスを選択することで、プロキシベンチマークを構築するフレームワークである。
14のモデル、4つのエージェントベンチマーク、19の非エージェントベンチマークによる実験では、PACE-Benchがエージェントスコアを予測し、LOOCVは絶対誤差(MAE)を4%以下、スピアマン相関は0.80以上、ペアワイズモデルの精度は85%で、いずれもエージェント評価コストの1%以下である。
論文 参考訳(メタデータ) (2026-07-02T10:59:03Z) - The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? [80.24951682268332]
本稿では,自律エージェント開発のためのフロンティアモデルのキャパシティをテストするための評価フレームワークであるMeta-Agent Challenge(MAC)を紹介する。
評価の整合性を確保するため、このフレームワークは報奨ハッキングに対する多層防御によって確保される。
メタエージェントは人間工学的な基本方針とほとんど一致せず、その一部はプロプライエタリなフロンティアモデルに支配されている。
論文 参考訳(メタデータ) (2026-06-03T04:58:17Z) - Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis [3.3237915628874632]
効果的なエージェント評価は、会話の質、効率性、およびエージェントエラーの体系的診断を取り入れて、正確性のみに留まらないと論じる。
エージェントの旋回効率と中間進捗を両立させる新しい指標を提案する。
TEDフレームワークは、モデルとユーザの専門知識レベルをまたいだエージェントパフォーマンスに関する新たな洞察を明らかにします。
論文 参考訳(メタデータ) (2026-03-16T16:14:28Z) - Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation [87.47155146067962]
数百のタスクで並列評価をオーケストレーションする,標準化された評価ハーネスを提供する。
モデル、足場、ベンチマークにまたがる3次元解析を行う。
私たちの分析では、ほとんどのランで精度を低下させる高い推論努力など、驚くべき洞察が示されています。
論文 参考訳(メタデータ) (2025-10-13T22:22:28Z) - How can we assess human-agent interactions? Case studies in software agent design [52.953425368394306]
我々は,人間とエージェントの相互作用の厳密な評価に向けて,二つの大きな一歩を踏み出した。
エージェント設計のより効率的な人間中心評価のためのフレームワークであるPULSEを提案する。
私たちは、オープンソースのソフトウェアエージェントOpenHandsを中心に構築された大規模なWebプラットフォームにフレームワークをデプロイします。
論文 参考訳(メタデータ) (2025-10-10T19:04:28Z) - Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation [4.08768677009363]
本稿では,タスク領域に依存しないエージェントタスク完了を評価するための,汎用的でモジュール化されたフレームワークを提案する。
GAIAとBigCodeBenchの2つのベンチマークでMagentic-One Actor Agentを評価することで、我々のフレームワークを検証する。
我々の審査員は、人間の評価と密接に一致したタスクの成功を予測し、それぞれ4.76%と10.52%のアライメント精度を達成した。
論文 参考訳(メタデータ) (2025-08-07T15:39:48Z) - Establishing Best Practices for Building Rigorous Agentic Benchmarks [94.69724201080155]
多くのエージェントベンチマークがタスク設定や報酬設計に問題があることを示す。
このような問題は、エージェントのパフォーマンスを最大100%相対的に過小評価することにつながる可能性がある。
我々はベンチマーク構築経験から要約したガイドラインの集合であるAgentic Benchmark Checklist (ABC)を紹介した。
論文 参考訳(メタデータ) (2025-07-03T17:35:31Z) - AutoPenBench: Benchmarking Generative Agents for Penetration Testing [42.681170697805726]
本稿では,自動貫入試験における生成エージェント評価のためのオープンベンチマークであるAutoPenBenchを紹介する。
エージェントが攻撃しなければならない脆弱性のあるシステムを表す33のタスクを含む包括的フレームワークを提案する。
完全自律型と半自律型という2つのエージェントアーキテクチャをテストすることで,AutoPenBenchのメリットを示す。
論文 参考訳(メタデータ) (2024-10-04T08:24:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。