論文の概要: The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
- arxiv url: http://arxiv.org/abs/2610.08026v1
- Date: Tue, 06 Oct 2026 09:19:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 02:58:29.917098
- Title: The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
- Title(参考訳): 幻覚検出ベンチマークにおけるラベル付け問題:実証的評価
- Abstract要約: 我々は,大規模言語モデル (LLM) がオープンドメイン質問応答データセットで幻覚することを検出する方法のベンチマークを行った。
この評価設定は、基準忠実性と事実正当性という2つの基準の方法論的曖昧さを生成する。
3つのQAデータセットと3つのジェネレータモデルにまたがる900人のラベル付き質問応答ペアを用いて,この潜在的な基準ミスマッチについて検討した。
- 参考スコア(独自算出の注目度): 0.2799896314754614
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
- Abstract(参考訳): 近年,大規模言語モデル (LLM) の幻覚を検知する手法が開発されている。
これらの手法は、しばしば、質問と対応する短い参照回答を含むオープンドメイン質問応答(QA)データセットでベンチマークされる。
まず、LLMを使用してQAデータセット内の質問に対する回答を生成する。
次に、データセットの参照回答と比較することにより、これらの回答を幻覚的か否かをラベル付けするために、いくつかの自動ラベリング戦略が使用される。
この評価設定は、参照忠実性(回答が参照によって完全に支持されているかどうか)と事実正しさ(回答が矛盾や事実的虚偽の特定の主張を含まないかどうか)の2つの基準の方法論的曖昧さを生成する。
実際には、意図した目標が後者であっても、自動ラベリングは前者の基準を適用することができる。
3つのQAデータセットと3つのジェネレータモデルにまたがる900の人間ラベル付き質問応答ペアを用いて,この潜在的な基準ミスマッチについて検討した。
我々は,レキシカル類似度指標,基準付きNLIベースライン,および制御されたプロンプト変種の下でのLLM審査員7名を自動ラベラとして評価した。
実験の結果,自動ラベル付け戦略と,これらのラベルと人的アノテーションの相違が明らかとなった。
多くの戦略は、強い方向性の誤りバイアスを示しており、ほとんどの裁判官とジェネレーターのペアは、忠実志向のプロンプトを事実的正確性に置き換えることで、人間のアノテーションとの合意を改善し、偽陽性の優位性を低下させ、自動幻覚ラベルがターゲットの基準の指定方法に強く依存していることを示す。
したがって、ラベルソースの選択は、ベンチマーク設計の基本的な部分と見なされるべきであり、明示的で、検証され、ベンチマークの目標に合致する。
関連論文リスト
- How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation [38.482877747109946]
本研究では,8つのクラスに有意な回答を割り当てる意味的正当性分類を導入し,有意な内容によって汚染されたものから有意な回答を分離する。
次に、広く使用されているQAデータセットにまたがる8.8kサンプルベンチマークであるCAP-Correctnessと、質問応答ペアを宣言文に変換する11kサンプルデータセットであるCAP-Statementsをリリースする。
第三にCAP(Context-Aware Precision)は、双方向NLIを用いて質問条件文をスコアする参照ベースのメトリクスである。
論文 参考訳(メタデータ) (2026-09-01T15:06:39Z) - Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts [41.162545164426085]
大規模言語モデル(LLM)を用いた文脈におけるラベル検証について検討する。
主観的ラベル補正のためのLiaHR(Label-in-a-Haystack Rectification)フレームワークを提案する。
このアプローチは、信号と雑音の比率を高めるために、アノテーションパイプラインに統合することができる。
論文 参考訳(メタデータ) (2025-05-22T18:55:22Z) - Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3 [5.478764356647438]
そこで本稿では,大規模言語モデル (LLM) を付加する代替手法を提案する。
基準レベルのグレードを関連ラベルに集約する様々な方法を検討する。
2024年夏に発生した LLMJudge Challenge のデータをもとに,我々のアプローチを実証的に評価する。
論文 参考訳(メタデータ) (2024-10-17T21:37:08Z) - Localizing Factual Inconsistencies in Attributable Text Generation [74.11403803488643]
本稿では,帰属可能なテキスト生成における事実の不整合をローカライズするための新しい形式であるQASemConsistencyを紹介する。
QASemConsistencyは、人間の判断とよく相関する事実整合性スコアを得られることを示す。
論文 参考訳(メタデータ) (2024-10-09T22:53:48Z) - Appeal: Allow Mislabeled Samples the Chance to be Rectified in Partial Label Learning [55.4510979153023]
部分ラベル学習(PLL)では、各インスタンスは候補ラベルのセットに関連付けられ、そのうち1つだけが接地真実である。
誤記されたサンプルの「アペアル」を支援するため,最初の魅力に基づくフレームワークを提案する。
論文 参考訳(メタデータ) (2023-12-18T09:09:52Z) - Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object
Detection with Repeated Labels [6.872072177648135]
そこで本研究では,基礎的真理推定手法に適合する新しい局所化アルゴリズムを提案する。
また,本アルゴリズムは,TexBiGデータセット上でのトレーニングにおいて,優れた性能を示す。
論文 参考訳(メタデータ) (2023-09-18T13:08:44Z) - Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection [98.66771688028426]
本研究では,一段階検出器のためのAmbiguity-Resistant Semi-supervised Learning (ARSL)を提案する。
擬似ラベルの分類とローカライズ品質を定量化するために,JCE(Joint-Confidence Estimation)を提案する。
ARSLは、曖昧さを効果的に軽減し、MS COCOおよびPASCALVOC上で最先端のSSOD性能を達成する。
論文 参考訳(メタデータ) (2023-03-27T07:46:58Z) - MatchGAN: A Self-Supervised Semi-Supervised Conditional Generative
Adversarial Network [51.84251358009803]
本稿では,条件付き生成逆数ネットワーク(GAN)に対する,半教師付き環境下での自己教師型学習手法を提案する。
利用可能な数少ないラベル付きサンプルのラベル空間から無作為なラベルをサンプリングして拡張を行う。
本手法は,ベースラインのトレーニングに使用したラベル付きサンプルの20%に過ぎません。
論文 参考訳(メタデータ) (2020-06-11T17:14:55Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。