論文の概要: CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
- arxiv url: http://arxiv.org/abs/2608.18554v1
- Date: Wed, 19 Aug 2026 05:22:20 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-20 20:13:55.286544
- Title: CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
- Title(参考訳): CentaurBench: 実世界の作業タスクの自動化に対するLLM機能のベンチマーク
- Authors: Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, Abhishek Nagaraj,
- Abstract要約: ほとんどのベンチマークでは、作業タスクの自動化に関するモデルをランク付けしている。
しかし、実際には、モデルは他の(人間またはLLM)エージェントを支援するためにしばしば使用される。
我々は、他のエージェントのパフォーマンスを自動化および拡張するモデルの能力を評価する統一的なフレームワークを導入する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
- Abstract(参考訳): ほとんどのLCMベンチマークは、作業タスクの自動化に関するモデルをランク付けしている。
しかし、実際には、モデルは他の(人間またはLLM)エージェントを支援するためにしばしば使用される。
したがって、モデル選択を駆動する問題は、どのモデルが最高の出力を生成するかだけではなく、どのモデルが他の(ウィーカー)エージェントの作業を改善するかである。
我々は、他のエージェントのパフォーマンスを自動化および拡張するモデルの能力を評価する統一的なフレームワークを導入する。
7つの経済的基盤を持つ実世界のタスクに対して、アシスタントモデルは、標準化された低容量労働者モデルのための支援テキストを書き、成果物を生成する。
自動化モードでは、アシスタントは出力を直接生成する。
アウトプットは10回のランで複製されるタスク固有のルーリックとLCMの審査パネルで盲対比較によって得られる。
2つの体制のランクはわずかに相関しているだけであり、自動化の勝者は7つのタスクのうち5つで増員を失う。
援助は確実に肯定的ではない。
不明な作業員は、3つのタスクですべてのアシスト条件を上回り、1つのモデルのガイダンスだけが平均的なガイダンスを上回りません。
これらの結果は、自動化能力が補助品質の不完全なプロキシであり、人間-AIやマルチエージェントシステムでの役割に応じてモデルを評価するベンチマークを動機付けていることを示唆している。
関連論文リスト
- Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job? [0.0]
SelfScoreは、ヘルプデスクとプロのコンサルティングタスクにおけるLLM(Large Language Model)の自動エージェントのパフォーマンスを評価するために設計されたベンチマークである。
このベンチマークは、問題の複雑さと応答の助け、スコアリングシステムにおける透明性と単純さの確保に関するエージェントを評価する。
この研究は、特にAI技術が優れている地域では、労働者の移動の可能性への懸念を提起している。
論文 参考訳(メタデータ) (2024-10-05T14:37:35Z) - AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning [54.47116888545878]
AutoActはQAのための自動エージェント学習フレームワークである。
大規模アノテートデータやクローズドソースモデルからの合成計画軌道は依存していない。
論文 参考訳(メタデータ) (2024-01-10T16:57:24Z) - TaskBench: Benchmarking Large Language Models for Task Automation [82.2932794189585]
タスク自動化における大規模言語モデル(LLM)の機能を評価するためのフレームワークであるTaskBenchを紹介する。
具体的には、タスクの分解、ツールの選択、パラメータ予測を評価する。
提案手法は, 自動構築と厳密な人的検証を組み合わせることで, 人的評価との整合性を確保する。
論文 参考訳(メタデータ) (2023-11-30T18:02:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。