論文の概要: Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
- arxiv url: http://arxiv.org/abs/2610.00296v1
- Date: Sat, 26 Sep 2026 12:56:58 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.596589
- Title: Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
- Title(参考訳): 正しいことだけじゃない:LLM推論におけるToken-Levelの確実性を再考
- Abstract要約: トークンレベルの確実性は、LLMトレーニングや推論において、正確性のためのプロキシとして広く使用されている。
モデルおよびタスク間での制御された経験的評価により、確実性の予測能力を評価する。
- 参考スコア(独自算出の注目度): 25.24545341768214
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answer correctly and distinguishing correct from incorrect responses to the same question. In our experiments, certainty is generally better at identifying questions a model is likely to answer correctly than at distinguishing correct from incorrect responses to the same question. Certainty also varies systematically across token types and positions within words, reflecting local properties of words and text form. Information about question difficulty appears early in generation, while the weaker information about answer correctness is more concentrated near the end. These findings show that the information certainty provides for decisions depends on the prediction target, the model, the certainty metric, and which token positions in the response are included in aggregation. We further demonstrate the practical value of these findings for test-time compute. We allocate the number of responses using certainty early in generation and weight answer votes using certainty near the end of each response. Compared with a fixed-sampling majority-voting baseline, this approach increases overall accuracy from 78.71\% to 79.54\% while reducing generated-token cost by 82.4\%.
- Abstract(参考訳): トークンレベルの確実性は、LLMトレーニングや推論において、正確性のためのプロキシとして広く使用されている。
しかし、確実性に基づく手法の性能は、確実性スコアの情報とそれらのスコアの使用方法の両方に依存する。
そこで我々は,モデルとタスク間での制御された経験的評価を通じて,確実性の予測能力を直接評価する。
モデルが正しい解答をする確率が高い質問を識別し、同じ質問に対する誤った応答と正しい解答を区別する。
我々の実験では、確実性は、モデルが正しい答えをする可能性のある質問を特定するのに、同じ質問に対する間違った反応と正しい答えを区別するよりも、一般的に優れている。
確実性はまた、トークンの種類や単語内の位置によって体系的に変化し、単語やテキスト形式の局所的な特性を反映している。
質問の難易度に関する情報は生成初期に出現するが、回答の正しさに関する弱い情報は終末近くに集中している。
これらの結果から,情報確実性は予測対象,モデル,確実度指標,応答中のトークン位置などに依存することが明らかとなった。
テスト時間計算におけるこれらの知見の実用的価値をさらに示す。
各応答の終端付近で、早期生成と重み付けの回答投票を用いて、応答数を割り当てる。
固定サンプリングの多数決投票ベースラインと比較して、このアプローチは全体の精度を78.71\%から79.54\%に引き上げ、生成トーケンコストを82.4\%削減する。
関連論文リスト
- Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration [3.450547277166974]
LLMの信頼性校正は、トークン確率スコアと言語的信頼の2つの信号を比較することで評価されることが多い。
我々は、動詞化-vs-token比較を定義する測度軸を変化させる。
両信頼性信号はプロトコルに依存した行動測定として扱うべきである。
論文 参考訳(メタデータ) (2026-05-26T23:03:38Z) - Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy [41.17236645999555]
本研究では,現実的な質問応答や数学的推論タスクにおいて,意味保存的言い回しの下でモデル解がどう変化するかを検討する。
4つのベンチマークと13のモデルで、モデル出力はプロンプトの正確な表現に依存することが多い。
多くの質問に対して、モデルはフレーズによって正しい答えと間違った答えを交互に行い、ミスマッチ率は23%以上に達する。
論文 参考訳(メタデータ) (2026-05-18T16:45:13Z) - Reference-Free Rating of LLM Responses via Latent Information [53.463883683503106]
本研究では,判断モデルに対して,自由テキスト応答にQuattスケールのスコアを割り当てるよう依頼する一般的な実践について検討する。
次に、内部モデル信号からスカラー評価を導出する潜在裁判官を提案し、評価する。
ペアとシングルレーティングのベンチマークの幅広いスイートの中で、潜在メソッドは標準のプロンプトにマッチするか、超えている。
論文 参考訳(メタデータ) (2025-09-29T12:15:52Z) - Probabilistic Modeling of Disparity Uncertainty for Robust and Efficient Stereo Matching [61.73532883992135]
本稿では,新しい不確実性を考慮したステレオマッチングフレームワークを提案する。
我々はベイズリスクを不確実性の測定として採用し、データを別々に見積もり、不確実性をモデル化する。
論文 参考訳(メタデータ) (2024-12-24T23:28:20Z) - Rethinking LLM Uncertainty: A Multi-Agent Approach to Estimating Black-Box Model Uncertainty [47.95943057892318]
ブラックボックスLSMの不確実性の定量化は、信頼性の高い応答とスケーラブルな監視に不可欠である。
本研究では,不確実性推定にマルチエージェント相互作用を用いた新しい理論的基礎手法であるDiverseAgentEntropyを紹介する。
論文 参考訳(メタデータ) (2024-12-12T18:52:40Z) - MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty [10.154013836043816]
データ不確実性の存在下での過去の不確実性定量化手法について検討する。
以上の結果から,従来の手法はシングル・アンサー・セッティングに比べて比較的困難であったことが示唆された。
我々は,データ不確実性が存在する場合でも,エントロピーと一貫性に基づく手法がモデル不確実性を効果的に推定することを示した。
論文 参考訳(メタデータ) (2024-08-13T11:17:31Z) - Evaluating language models as risk scores [23.779329697527054]
質問応答 LLM を用いてリスクスコアを生成するソフトウェアパッケージである folktexts を紹介する。
提案した5つのベンチマークタスクにまたがって17の最近のLCMを評価した。
複数選択質問応答によるゼロショットリスクスコアは高い予測信号を持つが、広く誤校正されている。
論文 参考訳(メタデータ) (2024-07-19T18:13:37Z) - Uncertainty is Fragile: Manipulating Uncertainty in Large Language Models [79.76293901420146]
大規模言語モデル(LLM)は、出力の信頼性が不可欠である様々な高い領域で採用されている。
本研究では,不確実性推定の脆弱性を調査し,攻撃の可能性を探る。
攻撃者がLSMにバックドアを埋め込むことができ、入力中の特定のトリガーによって起動されると、最終的な出力に影響を与えることなくモデルの不確実性を操作できることを示す。
論文 参考訳(メタデータ) (2024-07-15T23:41:11Z) - Improving the Reliability of Large Language Models by Leveraging
Uncertainty-Aware In-Context Learning [76.98542249776257]
大規模言語モデルはしばしば「ハロシン化」の課題に直面している
本研究では,不確実性に応答してモデルが出力を拡張あるいは拒否することを可能にする,不確実性を考慮したコンテキスト内学習フレームワークを提案する。
論文 参考訳(メタデータ) (2023-10-07T12:06:53Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。