論文の概要: Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
- arxiv url: http://arxiv.org/abs/2607.17044v1
- Date: Sun, 19 Jul 2026 03:20:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.343731
- Title: Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
- Title(参考訳): エージェントの信頼性はどこから来るのか? - 検証ループ, スペシャリストモデル, およびプロダクションエージェントにおけるスキャッディングのクロスベンチマーク分解-
- Authors: Arunabh Dastidar, the Leni Team,
- Abstract要約: アーキテクチャが検証ループをインストールする1つのプロダクションシステムについて検討する。
ループの端から端までを計測し、経験的検証器混同行列を生成する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Multi-step enterprise agent tasks fail in a characteristic way: single-pass inference has no checkpoint between deciding an answer and committing to it. We study one production system (Leni) whose architecture installs such checkpoints: verification loops (execute, observe, compare, correct) staffed by lightweight task-specialized post-trained models. We evaluate the unmodified production configuration on three public benchmarks stressing distinct failure modes: SpreadsheetBench Verified (silent computation error), BullshitBench v2 (premise confabulation), and the GAIA validation split (cascade error over long tool chains). The full system improves over its frontier base model by +11.0 percentage points on SpreadsheetBench (91.25% vs 80.25%, n=400, p<0.001), +7 to +10 percentage points on BullshitBench (98% vs 91%, n=100), and roughly +15 points on GAIA validation (75.2% pass@1, n=165; 83.0% best-of-k). Our central contribution is a decomposition of that uplift: most of it comes from scaffolding, routing, and specialist models rather than from the verification step itself, whose isolated contribution is small (+1.5 points) but concentrated at the top of the score distribution, where it converts otherwise-failing tasks. We instrument the loop end-to-end, yielding an empirical verifier confusion matrix (catch rate about 0.20, fix rate 0.75, no false-alarm regressions) that grounds a compounding-reliability model. Specialist-swap ablations suggest that the loop's value depends on who observes it: replacing the small trained verifier with the generating frontier model eliminates most rescues. A valid-premise control shows zero over-rejections in 100 expert-level questions.
- Abstract(参考訳): マルチステップのエンタープライズエージェントタスクは、特徴的な方法で失敗する。
このようなチェックポイントをアーキテクチャがインストールする1つのプロダクションシステム(Leni)について検討する。
SpreadsheetBench Verified(サイレントな計算誤差)、BullshitBench v2(前提の折り畳み)、GAIA Validation split(長いツールチェーン上のカスケードエラー)である。
SpreadsheetBench (91.25% vs 80.25%, n=400, p<0.001), +7 to +10% on BullshitBench (98% vs 91%, n=100), and almost +15% on GAIA validation (75.2% pass@1, n=165; 83.0% best-of-k)。
私たちの中心的なコントリビューションは、そのアップリフトの分解です。大部分が、検証ステップ自体からではなく、足場、ルーティング、スペシャリストモデルによるもので、独立したコントリビューションは小さい(+1.5ポイント)が、スコア分布の上部に集中しており、そうでないタスクを変換します。
ループの終端から終端までを計測し,複合信頼性モデルに基づく経験的検証器混同行列(キャッチレート約0.20,固定率0.75,偽アラーム回帰なし)を導出する。
専門家とスワップのアブレーションは、ループの価値は誰がそれを観察するかに依存し、小さな訓練された検証器を生成フロンティアモデルに置き換えることで、ほとんどの救助を排除していることを示唆している。
有効な前提制御は、専門家レベルの100の質問でオーバーリジェクションをゼロにする。
関連論文リスト
- Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI [2.2616066835313253]
7つのカテゴリと3つの難易度で86のPicoCTF課題を実行し、3つのツールアクセスレジームと3つのモデル/クライアント構成で実行します。
既存のツール,エージェント・ビヘイビアの変更,11の新たな機能ツールに修正を適用し,これまで未経験だったトライアルを再実行します。
全体の解決率は55.4%から72.0%に上昇し、全ての構成が大幅に改善される。
論文 参考訳(メタデータ) (2026-07-03T02:20:04Z) - Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction [55.11480729304395]
PointArena 2026は77.2%の精度でベンチマークで2位である。
ap proachは3つの障害モードをターゲットにしている。第一に、エージェント駆動のシンセシスは大きなセマンティクスとアンカー相対的な候補プールを構築する。
次に、determinis tic steerable-dataパイプラインは、認証された10,000サンプルのメインセットと、マスク、テンプレート、パス検証を使用するリザーブサンプルを生成する。
論文 参考訳(メタデータ) (2026-06-29T06:39:03Z) - SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution [6.908637308550535]
SEVAは、エビデンスアライメント、ステップバイステップの推論チェーン、キャリブレーションされた信頼度、6カテゴリのエラー診断を発行する構造化検証エージェントである。
ClearFacts では、SEVA-3B は GPT-4o-mini (69.0 vs. 69.8 F1) と一致し、よりリッチで監査可能な出力を生成する。
論文 参考訳(メタデータ) (2026-06-29T02:37:13Z) - Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning [40.342912574072024]
大規模言語モデルは、旅行計画やコードソリューションのような構造化されたアウトプットを生成する。
個々の推論ステップは正しく見えるが、アウトプット全体が予算に違反したり、テストケースに失敗したり、あるいは以前の推論に矛盾することがある。
構造化LCM出力の検証のための決定論的解析制約付き学習品質スコアラを提案する。
論文 参考訳(メタデータ) (2026-05-15T17:08:27Z) - CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing [5.661334639541121]
CRANEは、シンキング・インストラクトデルタを、インストラクトバックボーンの候補推論編集のプールとして扱う、トレーニング不要なパラメータ編集手法である。
ペア化されたインストラクトとシンキングのチェックポイントを組み合わせることで、CRANEはどちらのモデルよりも強力なゲインを提供する。
論文 参考訳(メタデータ) (2026-05-13T20:09:35Z) - Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models [94.68358825189738]
本稿では,予測精度と推論品質を協調的に最適化する検証済み領域の学習後フレームワークを提案する。
我々は,エンジン信号に対して推論ステップを確定的に検証できる制御テストベッドであるチェスのVPSを評価する。
VPSは、推論品質を著しく向上させながら精度を保ち、勝利率エラーを最大30%削減し、一貫性をほぼ飽和状態に回復する。
論文 参考訳(メタデータ) (2026-04-03T15:19:46Z) - HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam [63.84155758655084]
HumanityのLast Exam (HLE)は、フロンティアの大規模言語モデルを評価するために広く使われているベンチマークである。
HLE-Verifiedは,透過的検証プロトコルときめ細かい誤り分類法を備えたHLEの検証および改訂版である。
我々は,HLEとHLE-Verifiedの7つの最先端言語モデルを評価し,平均7~10ポイントの絶対精度を観測した。
論文 参考訳(メタデータ) (2026-02-15T02:50:15Z) - SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents [52.20768003832476]
我々は$$-Bench (Airline/Retail) および SWE-Bench Verified 上での実行トレースを分析する。
成功を失敗に戻すための、先進的な逸脱、最初期の行動、レベル分岐を形式化する。
モデルに依存しない,勾配のない,テスト時のセーフガードである cm を導入します。
論文 参考訳(メタデータ) (2025-11-26T01:28:22Z) - Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems [20.846301581161978]
マルチエージェントシステムにおける障害帰属は、批判的だが未解決の課題である。
現在の手法では、これを長い会話ログ上のパターン認識タスクとして扱う。
A2P Scaffoldingは、パターン認識から構造化因果推論タスクへの障害帰属を変換する。
論文 参考訳(メタデータ) (2025-09-12T16:51:15Z) - TACRED Revisited: A Thorough Evaluation of the TACRED Relation
Extraction Task [80.38130122127882]
TACREDはリレーショナル抽出(RE)において最も大きく、最も広く使われているクラウドソースデータセットの1つである
パフォーマンスの天井に到達したのか、改善の余地はあるのか?
ラベルエラーは絶対F1テストエラーの8%を占めており、例の50%以上を可逆化する必要がある。
論文 参考訳(メタデータ) (2020-04-30T15:07:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。