論文の概要: Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents
- arxiv url: http://arxiv.org/abs/2606.10315v1
- Date: Tue, 09 Jun 2026 02:11:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-10 15:40:58.258726
- Title: Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents
- Title(参考訳): LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents (特集:LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents)
- Authors: Sawyer Zhang, Alexander Wang, Sophie Lei,
- Abstract要約: デプロイされた多ターン食品・飲料注文エージェントについて検討し,実際の品質問題の数を測定した。
私たちの盲点分類は、失敗はランダムではなく構造化されていることを示している。
プロダクションマルチターンエージェントでは、自動判断はリグレッションフロアであり、人間のレビューの代わりにはならない。
- 参考スコア(独自算出の注目度): 45.148328075418156
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: LLM-as-judge is the default instrument for evaluating conversational agents, yet its reliability is almost always reported as agreement with human ratings, not recall of real defects. We study a deployed multi-turn food-and-beverage ordering agent and measure how many genuine quality problems its built-in LLM judge catches, using exhaustive human transcript review as ground truth. Across three batches the judge surfaces well under a quarter of human-confirmed systematic problems -- 2 of 9 patterns (22%) in one batch, and its operational gate flagged zero of 100 rounds in a batch where humans confirmed 23 distinct defects and 7 new cross-cutting patterns. Our blind-spot taxonomy shows the failure is structured, not random: the judge catches turn-local issues (a fabricated statistic, a wrong language) but misses cross-turn state issues (confirm-gate lockout, cart hallucination, escalation lockout, stale referents). The mechanism: the scoring rubric exposes only three coarse axes (intent, brand-voice, personalization) and has no category for the behavioural dimensions -- state-tracking, guardrails, recovery -- where most defects cluster. The failure is routing, not perception: 113 of 114 rounds whose raw judge note describes a confirm-gate or cart-state defect are scored "brand voice", and none reach an operational failure -- the gate is wired to hangs and hard assertions, not the rubric -- so the 0% is a routing-and-wiring failure, not blindness. The consequence for prevalence estimation is sharp: when the apparent defect rate is zero the Rogan-Gladen correction degenerates -- no signal can recover the true rate -- while where the gate reports a nonzero rate the same estimator implies a 3-6x undercount under our measured sensitivity. For production multi-turn agents, automated judging is a regression floor, not a substitute for human review.
- Abstract(参考訳): LLM-as-judgeは会話エージェントを評価するためのデフォルトの手段であるが、その信頼性はほとんど常に人間の評価と一致していると報告されている。
我々は,多ターン食品・飲料注文エージェントを配備し,そのLCM判定器がキャッチする真の品質問題の数を測定し,徹底的な人文書評を根拠として検討した。
3つのバッチにわたって、審査員は、人間確認された体系的な問題の4分の1(9つのパターンのうち22%)を1バッチで覆い、その運用ゲートは、23の異なる欠陥と7つの新しい横断的なパターンが確認されたバッチで100ラウンドのフラグを立てた。
私たちの盲点分類は、失敗は構造化されており、ランダムではないことを示している。裁判官は、ターンローカルな問題(製造された統計学、間違った言語)をキャッチするが、クロスターン状態の問題(確認ゲートロックアウト、カートの幻覚、エスカレーションロックアウト、古い参照)を見逃す。
スコアリングルーブリックは3つの粗い軸(インテント、ブランドボイス、パーソナライゼーション)のみを露出し、ほとんどの欠陥がクラスタ化されている状態追跡、ガードレール、リカバリといった振る舞いの次元のカテゴリを持たない。
確認ゲートやカート状態の欠陥を記載した114ラウンドのうち、113ラウンドは「ブランドボイス」と評価され、運用上の障害には到達せず、ゲートはハングやハードアサーションに配線されており、ルーリックではないため、0%はルーティングと配線の失敗であり、盲目ではない。
明らかな欠陥率がゼロである場合、ロガン・グラデン補正は縮退する -- 真の速度を回復できない -- 一方で、ゲートはゼロではない速度を報告し、同じ推定器は測定感度の下で3~6倍のアンダーカウントを示す。
プロダクションマルチターンエージェントでは、自動判断はリグレッションフロアであり、人間のレビューの代わりにはならない。
関連論文リスト
- Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench [0.0]
AgentProp-Benchは4つのドメインに2300のトレースを持つ2,000タスクのベンチマークである。
我々は、判断信頼性を定量化し、エラーの伝播を特徴づけ、実行時の緩和を評価する。
すべてのコード、データ、トレース、および人間のラベルはhttps://github.com/bhaskargurram-ai/agenthallu-bench.orgで公開されている。
論文 参考訳(メタデータ) (2026-04-17T21:15:35Z) - Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications [51.56484100374058]
我々は,エビデンスに基づくリリース決定を伴う品質ゲートを導入する自動自己テストフレームワークを提案する。
内部展開型多エージェント対話型AIシステムの縦型ケーススタディにより,本フレームワークの評価を行った。
論文 参考訳(メタデータ) (2026-03-13T20:44:15Z) - A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness [57.510025257780306]
既存の検証プロトコルは、レッドチーム固有の分散シフトを考慮できないことを示す。
我々は、より一貫して判断可能な振る舞いのベンチマークであるReliableBenchと、判断失敗を公開するために設計されたデータセットであるJiceStressTestを提案する。
論文 参考訳(メタデータ) (2026-02-04T15:13:35Z) - Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior [0.0]
審査員は一貫性があるが、互いに一致していない。
評価は3,240件を超え、中間合意はほぼゼロに近い。
審査員の平均得点は、審査員の実際の値に該当しない合成判定を生成する。
論文 参考訳(メタデータ) (2026-01-08T17:02:22Z) - CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal [84.71254539482369]
検証可能な報酬を伴うグループ相対的強化学習(RLVR)は、しばしば、すでに失敗している最も情報に富むデータを浪費する。
エラーを監督するマルチモーダル推論のための,障害中心のポストトレーニングフレームワークであるCAREを提案する。
CAREは正確さを改善し、スムーズさをトレーニングすると同時に、障害からの学習信号のシェアを明示的に増やします。
論文 参考訳(メタデータ) (2025-12-22T16:34:21Z) - TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them [58.04324690859212]
自動評価器(LLM-as-a-judge)としての大規模言語モデル(LLM)は、現在の評価フレームワークにおいて重大な矛盾を明らかにしている。
スコア比較不整合とペアワイズ・トランジティビティ不整合という2つの基本的不整合を同定する。
我々は2つの重要なイノベーションを通じてこれらの制限に対処する確率的フレームワークであるTrustJudgeを提案する。
論文 参考訳(メタデータ) (2025-09-25T13:04:29Z) - When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity [21.192000569821943]
我々は、厳密な目標と検証可能な構成がなければ、ベンチマークのランキングは、ほぼノイズの多い高信頼度ランキングを生成することができると論じる。
本稿では,Arena-Hard Autoが使用するELOスタイルのアグリゲーションが崩壊し,真のランキングの不確かさをマスクすることを示す。
我々の結果は、妥当性を損なう設計上の失敗を強調し、より良いスコープで信頼性に配慮したベンチマークを構築するための実用的な原則を提供する。
論文 参考訳(メタデータ) (2025-09-24T16:26:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。