論文の概要: Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses
- arxiv url: http://arxiv.org/abs/2609.03230v1
- Date: Thu, 03 Sep 2026 00:06:04 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-04 18:28:38.882272
- Title: Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses
- Title(参考訳): 2つの真実と嘘? 要求品質評価のための市販LCMのベンチマーク:パフォーマンス、偽アラーム、失敗
- Authors: Jannatul Shefa, Alejandro Salado, Paul Wach, Taylan G. Topcu,
- Abstract要約: 大規模言語モデル(LLM)は要求品質の評価を吸収することができる。
本研究では,要求品質評価のための既設LLM性能のベンチマーク解析を行った。
強い非対称な誤差プロファイルを定量化する: 全てのモデルと実行において、最高のパフォーマンスの人類学モデルは、専門家が特定した問題の47%の中央値を検出し、偽フレーガーは11%である。
- 参考スコア(独自算出の注目度): 39.22505505045623
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns. Because requirements are often written in natural language, recent advances in generative AI have raised expectations that large language models (LLMs) can absorb requirement quality assessment, a task otherwise slow and human expertise-intensive. Yet empirical evidence on whether LLMs can be trusted to do so remains scarce. This study presents the first benchmarking analysis of off-the-shelf LLM performance for requirement quality evaluation. Against an expert-derived ground truth built on INCOSE quality criteria, we evaluate ten models spanning two families (OpenAI and Anthropic) and five generations each, across one hundred independent runs, two requirement sets, and five sampling temperatures. Four contributions follow. First, we quantify a strongly asymmetric error profile: across all models and runs, the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%. Second, performance degrades significantly where SE judgment is required, as necessity and correctness issues are almost always missed. Third, generational progress is non-monotonic, so newer models cannot be assumed better. Fourth, this error behavior shifts only modestly and non-monotonically across sampling temperatures, indicating characteristic model deficiencies rather than inherent stochasticity. Off-the-shelf LLMs are therefore not yet trustworthy autonomous evaluators. Findings also warrant caution for Agentic AI developers: orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them. Their defensible near-term role is human-in-the-loop decision support.
- Abstract(参考訳): 要件エンジニアリング(RE)はシステムエンジニアリング(SE)の下流にあるものすべての品質を管理します。
要件はしばしば自然言語で記述されるため、生成AIの最近の進歩は、大きな言語モデル(LLM)が要求品質の評価を吸収できるという期待を高めている。
しかし、LSMを信頼できるかどうかに関する実証的な証拠は乏しい。
本研究では,要求品質評価のための既設LLM性能のベンチマーク解析を行った。
InCOSEの品質基準に基づく専門家による根拠的真理に対して,我々は,100の独立ラン,2つの要求セット,5つのサンプリング温度にまたがる2つの家族(OpenAI, Anthropic)と5世代にまたがる10のモデルを評価する。
以下の4つのコントリビューションがある。
第一に、強い非対称な誤差プロファイルを定量化する: 全てのモデルと実行において、最高のパフォーマンスの人類学モデルは、専門家が特定した問題の47%の中央値を検出し、偽フレーガーの11%を検知する。
第2に、SE判定が必要なところでパフォーマンスが著しく低下する。
第三に、世代別進行は単調ではないため、新しいモデルはより良く仮定することはできない。
第4に, この誤差の挙動は, サンプリング温度でわずかに, 非単調にのみ変化し, 固有確率ではなく, 特性モデル欠陥を示す。
したがって、市販のLCMは、まだ信頼できる自律評価器ではない。
これらのLLMモジュールを特殊なアーキテクチャでオーケストレーションすることは、修正するのではなく、これらの欠陥を複雑化するリスクがある。
彼らの防御可能な短期的役割は、人間によるループ内意思決定支援である。
関連論文リスト
- OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation [57.505743202759646]
OccuBenchは10の業界カテゴリと65の専門ドメインにわたる100の現実のプロフェッショナルタスクシナリオをカバーするベンチマークである。
我々のマルチエージェント合成パイプラインは, 可溶性, 校正困難, 文書基底の多様性を保証した評価インスタンスを自動生成する。
論文 参考訳(メタデータ) (2026-04-13T00:27:32Z) - How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling [43.64673843846063]
大規模言語モデル(LLM)は推論ベンチマークにおいて高いパフォーマンスを達成しているが、エンドツーエンドを必要とする現実世界の問題を解決する能力は未だ不明である。
本稿では、専門家が検証した基準を用いて、モデリング段階間でのLCM性能を評価する問題指向の段階評価フレームワークを提案する。
論文 参考訳(メタデータ) (2026-04-06T15:58:47Z) - Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning [32.32593439144886]
振舞い校正された強化学習により、小さなモデルは不確実な定量化においてフロンティアモデルを超えることができる。
当社のモデルでは,GPT-5の0.207を超える精度向上率(0.806)を挑戦的なドメイン内評価において達成している。
論文 参考訳(メタデータ) (2025-12-22T22:51:48Z) - Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark [0.0]
フレーム問題とシンボルグラウンド問題(英語版)は歴史的に、伝統的なシンボルAIシステムでは解決不可能と見なされてきた。
本研究では,現代のLSMがこれらの問題に対処するために必要な認知能力を持っているかを検討する。
論文 参考訳(メタデータ) (2025-06-09T16:12:47Z) - Statistical Runtime Verification for LLMs via Robustness Estimation [0.0]
ランタイムクリティカルなアプリケーションにLLM(Large Language Models)を安全にデプロイするためには、逆の堅牢性検証が不可欠である。
ブラックボックス配置環境におけるLCMのオンライン実行時ロバスト性モニタとしての可能性を評価するために,RoMA統計検証フレームワークを適応・拡張するケーススタディを提案する。
論文 参考訳(メタデータ) (2025-04-24T16:36:19Z) - LLM2: Let Large Language Models Harness System 2 Reasoning [65.89293674479907]
大規模言語モデル(LLM)は、無数のタスクにまたがって印象的な機能を示してきたが、時には望ましくない出力が得られる。
本稿では LLM とプロセスベースの検証器を組み合わせた新しいフレームワーク LLM2 を紹介する。
LLMs2は妥当な候補を生成するのに責任を持ち、検証者は望ましい出力と望ましくない出力を区別するためにタイムリーなプロセスベースのフィードバックを提供する。
論文 参考訳(メタデータ) (2024-12-29T06:32:36Z) - MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs [55.20845457594977]
大規模言語モデル(LLM)は、問題解決と意思決定の能力の向上を示している。
本稿ではメタ推論技術を必要とするプロセスベースのベンチマークMR-Benを提案する。
メタ推論のパラダイムは,システム2のスロー思考に特に適しています。
論文 参考訳(メタデータ) (2024-06-20T03:50:23Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。