論文の概要: External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
- arxiv url: http://arxiv.org/abs/2610.02066v1
- Date: Thu, 01 Oct 2026 17:06:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.31919
- Title: External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
- Title(参考訳): 大規模言語モデルにおけるクロスモデルスパンレベル幻覚検出の隠れ状態探索による外部オブザーバの発見
- Abstract要約: 粒度, 粒度, 空間レベルの幻覚検出のための内部隠れ状態フレームワークを提案する。
実験の結果,本手法は幻覚の発症を効果的に分離できることが判明した。
本稿では,あるモデルが他のモデルの生成によって引き起こされる内部表現を観察する,新しいクロスモデル検出フレームワークを提案する。
- 参考スコア(独自算出の注目度): 3.2228025627337864
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
- Abstract(参考訳): 大規模言語モデル (LLM) が基礎的推論エンジンとしての役割を担っているため、幻覚の傾向は依然として重大な脆弱性である。
最近の内部状態プローブは、遅い外部検索システムに代わる有望な代替手段を提供するが、トークン単位のバイナリ分類タスクへの幻覚検出を大幅に削減し、セマンティックドリフトの構造的、シーケンシャルな境界を捕捉することができない。
本稿では,細粒度でスパンレベルの幻覚検出のための内部隠れ状態フレームワークを提案する。
レイヤワイズアクティベーションパターンを検査することにより、LLM生成における正確な幻覚のオンセットと継続トークンの検出を試みる。
実験の結果, 極度のクラス不均衡にもかかわらず, ランダムなベースライン上での精度・リコールAUCの大幅な改善を実現し, 幻覚発症の分離に成功した。
最終的に、あるモデルが他のモデルの生成によって引き起こされる内部表現を観察する新しいクロスモデル検出フレームワークを提案する。
外部オブザーバは、観測者がより小さいモデルである場合を含む、自身の幻覚発症のジェネレータの自己検出を一致または超えることができ、自己検出がオンセット局所化の天井ではないことを示唆する。
関連論文リスト
- UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations [61.85916280854257]
LVLM(Large Vision-Language Models)は印象的な視覚的推論と対話機能を実現するが、視覚入力によるコンテンツサポートを幻覚させることが多い。
textbfUniProbeは軽量で統一された学習可能な検出器で、凍ったLVLMの不均一な計算トレースを1つの前方パスからモデル化する。
論文 参考訳(メタデータ) (2026-08-11T12:01:59Z) - RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration [0.2696472814555309]
本稿では,検出ヘッドを大規模言語モデルに統合する幻覚認識型微調整手法であるRAGognizerを紹介する。
RAGognizerは、生成時の幻覚率を大幅に低減しつつ、最先端のトークンレベルの幻覚検出を実現する。
論文 参考訳(メタデータ) (2026-04-17T11:07:32Z) - Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling [78.78822033285938]
VLM(Vision-Language Models)は視覚的理解に優れ、視覚幻覚に悩まされることが多い。
本研究では,幻覚を意識したトレーニングとオンザフライの自己検証を統合した統合フレームワークREVERSEを紹介する。
論文 参考訳(メタデータ) (2025-04-17T17:59:22Z) - Enhancing Hallucination Detection through Noise Injection [9.582929634879932]
大型言語モデル(LLM)は、幻覚として知られる、もっとも不正確な応答を生成する傾向にある。
ベイズ感覚のモデル不確実性を考慮し,検出精度を著しく向上できることを示す。
サンプリング中にモデルパラメータの適切なサブセット、あるいは等価に隠されたユニットアクティベーションを摂動する、非常に単純で効率的なアプローチを提案する。
論文 参考訳(メタデータ) (2025-02-06T06:02:20Z) - AutoHall: Automated Hallucination Dataset Generation for Large Language Models [56.92068213969036]
本稿では,AutoHallと呼ばれる既存のファクトチェックデータセットに基づいて,モデル固有の幻覚データセットを自動的に構築する手法を提案する。
また,自己コントラディションに基づくゼロリソース・ブラックボックス幻覚検出手法を提案する。
論文 参考訳(メタデータ) (2023-09-30T05:20:02Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。