Fugu-MT 論文翻訳(概要): What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

論文の概要: What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

arxiv url: http://arxiv.org/abs/2605.02038v1
Date: Sun, 03 May 2026 20:05:08 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-05 20:33:50.055056
Title: What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models
Title（参考訳）: 単発精度に欠けていること:言語モデルの多変量信頼性監査
Authors: Ranit Karmakar, Jayita Chatterjee,
Abstract要約: シングルプロンプト精度は、言語モデルをベンチマークする主要な方法であるが、重要な信頼性障害を見逃す可能性がある。 15モデルオープンウェイトコーパスの評価を行い,5つの分類と推論ベンチマークによる10のインストラクトモデルに着目した信頼性解析を行った。まず、評価設計は結論を根本的に変えることができる。第2に、信頼信号は脆弱である。MMLU-Proでは、各プライマリモデルは、その精度と同一行上のトークン確率信頼の両方よりもかなり高い信頼度を言語的に報告し、単一のプロンプト変種における単一のモデルに対して、動詞のパースレートが崩壊する可能性がある。
参考スコア（独自算出の注目度）: 0.0
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-weight corpus, with the main reliability analyses focused on 10 instruct models across five classification and reasoning benchmarks under five prompt variants each, measuring accuracy, token-probability calibration, verbal-confidence calibration, verbal parse rate, and prompt-perturbation spread for every (model x dataset x variant) cell. We find three broad results. First, evaluation design can materially change the conclusion. Switching Expected Calibration Error (ECE) token from a raw to a label-set-normalised definition changes per-cell calibration by a mean absolute 0.149. More strikingly, pairing a chain-of-thought prompt with a first-character evaluator on ARC-Challenge reduces apparent accuracy by 72-88% across all five primary models; two independent repair procedures recover 93.8% and 102.7% of the lost performance, indicating an evaluator-side rather than model-side failure. Second, confidence signals are fragile. On MMLU-Pro, every primary model verbally reports confidence substantially above both its accuracy and its token-probability confidence on the same rows, and verbal parse rate can collapse for a single model on a single prompt variant. Third, prompt robustness does not track parameter count reliably. Across 10 instruct models, the correlation between model size and prompt-perturbation spread ranges from -0.244 to 0.474 across benchmarks. Taken together, these results show that reliability conclusions for small language models depend not only on the model being evaluated, but also on the evaluation pipeline used to measure it. We argue that calibration definitions, evaluator logic, verbal parseability, and prompt robustness should be reported explicitly when making reliability claims.
Abstract（参考訳）: シングルプロンプト精度は、言語モデルをベンチマークする主要な方法であるが、重要な信頼性障害を見逃す可能性がある。提案する15モデルオープンウェイトコーパスは,5つの分類および推論ベンチマークを対象とし,評価精度,トークン確率キャリブレーション,動詞信頼度キャリブレーション,動詞パース率,および各セル(モデル x データセット x 変種)に対して,それぞれ10種類のインストラクトモデルに着目した信頼性解析を行った。 3つの大きな結果が得られます。第一に、評価設計は結論を大幅に変えることができる。期待校正誤り(ECE)トークンを生からラベルセット正規化定義に切り替えると、セルごとの校正は平均0.149で変化する。さらに印象的なことに、ARC-Challenge上の最初の文字評価器とチェーン・オブ・シークレットのプロンプトを組み合わせることで、5つのプライマリモデル全てで明らかな精度が72-88%低下し、2つの独立した修復手順が93.8%と102.7%の損失を回復した。第二に、信頼信号は脆弱である。 MMLU-Proでは、各一次モデルは、その精度と同一行上のトークン確率の信頼度を大きく上回る信頼度を言語的に報告し、単一のプロンプト変種における単一のモデルに対して、動詞のパースレートが崩壊する可能性がある。第三に、素早い堅牢性はパラメータの数を確実に追跡しない。 10のモデルに対して、モデルサイズと急激な摂動拡散の相関はベンチマークで-0.244から0.474の範囲である。これらの結果から,小型言語モデルに対する信頼性の結論は,評価対象モデルだけでなく,測定に用いる評価パイプラインにも依存することが示された。信頼性の主張を行う際には、校正定義、評価者論理、動詞のパーセビリティ、即時堅牢性を明示的に報告すべきである。

論文の概要: What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

関連論文リスト