論文の概要: Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
- arxiv url: http://arxiv.org/abs/2607.17883v1
- Date: Mon, 20 Jul 2026 12:34:44 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.614139
- Title: Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
- Title(参考訳): 構築によるゼロ幻覚:信頼できるエンタープライズAIのための幻覚を意識した階層化監視
- Abstract要約: 我々は、「ゼロ幻覚」はモデルが持つ性質ではなく、システムが強制する性質であると主張する。
本稿では,幻覚を最小限の障害モードではなく,持続可能な障害モードとして扱う保証アーキテクチャであるHALOを提案する。
- 参考スコア(独自算出の注目度): 39.17006317116773
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true. The common response is to wait for a model that does not hallucinate. We argue that this is the wrong target. Large language models are, by construction, capable of generating unsupported text, and no amount of scale removes the possibility; a faithfulness judge bolted onto a raw model catches some errors but still ships others, and even well-curated retrieval pipelines have been shown to fabricate citations. We reframe the goal: "zero hallucination" is not a property a model possesses but a property a system enforces. We present HALO (Hallucination-Aware Layered Oversight), an assurance architecture which treats hallucination as a containable failure mode rather than an eliminable one. HALO composes six layers of defense: grounded generation over retrieved, approved content; constrained, deterministic execution that bounds where the model can err; multi-signal verification that scores every output for groundedness and hallucination using both an LLM judge and evidence-based checks against the source text; calibrated abstention, so the system declines rather than guesses when grounding is insufficient; total traceability of every retrieval, tool call, and generation; and continuous oversight that detects drift, alerts on threshold breaches, and closes the loop by regenerating and statistically validating improved agents. We detail each layer, give particular attention to evidence-based confidence (which verifies extractions against the source document rather than trusting the model's self-reported certainty), and illustrate the architecture on a regulated claims-extraction workload
- Abstract(参考訳): 企業は信頼できないAIエージェントをデプロイしない。そして、不信の最も暗黙の理由は幻覚である。
一般的な反応は、幻覚のないモデルを待つことです。
これは間違ったターゲットだと我々は主張する。
大きな言語モデルは、建設によって、サポート対象のテキストを生成することができ、スケールの量ではその可能性を排除できない。
ゼロ幻覚(zero hallucination)"は、モデルが所有するプロパティではなく、システムが強制するプロパティです。
本稿では,幻覚を最小限の障害モードではなく,持続可能な障害モードとして扱う保証アーキテクチャであるHALO(Hallucination-Aware Layered Oversight)を提案する。
HALOは6つの防衛レイヤで構成されている: 検索、承認されたコンテンツに対する接地生成、モデルが発効可能な領域に縛られる制約付き決定的実行、LLMの判断と証拠に基づくソーステキストに対するチェックの両方を用いて、接地と幻覚の両方のアウトプットをスコアするマルチシグナル検証、接地が不十分なときの推測よりもシステムの調整、すべての検索、ツールコール、生成のトータルトレーサビリティ、ドリフト、しきい値違反の警告、改善されたエージェントの再生と統計的検証によってループを閉じる連続監視。
それぞれのレイヤを詳述し、エビデンスに基づく信頼性(モデルの自己報告された確実性を信頼するのではなく、ソース文書に対して抽出を検証する)に特に注意を払って、規制されたクレーム抽出のワークロードにアーキテクチャを記述します。
関連論文リスト
- The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents [0.0]
本稿では,ポストホックではなく構造的検証を行うためのエージェント機器を提案する。
組織ごとの書き込みエラー、レンダリングサイズ、あるいは塩漬けのカナリアエチョフロアが破られた場合、実行はすべて無効になる。
長期のエージェントが苦しむたびに、クリーンでシングル変数の結果が報告される。
論文 参考訳(メタデータ) (2026-08-04T15:10:37Z) - Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory [0.0]
ソース属性は、会話記憶の構造化方法に依存することを示す。
これは、AIシステムが自律的なマルチターンの役割を担い、彼らが知っていることを評価するのに十分ではないことを示唆している。
論文 参考訳(メタデータ) (2026-07-27T01:47:12Z) - Building Reliable Long-Form Generation via Hallucination Rejection Sampling [26.68294442366694]
大規模言語モデル(LLM)は、オープンエンドテキスト生成において顕著な進歩を遂げてきたが、誤ったあるいはサポートされていないコンテンツを幻覚させる傾向にある。
我々は、Segment-wise HAllucination Rejection Smpling (SHARS) という、新しい推論時幻覚緩和フレームワークを提案する。
SHARSは任意の幻覚検出器を使用して、生成中の幻覚セグメントを識別および拒絶し、忠実な内容が生成されるまで再サンプリングする。
論文 参考訳(メタデータ) (2026-06-02T13:26:17Z) - Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework [35.29785249891566]
大規模視覚言語モデル(LVLM)は、既存のオブジェクトの記述を生成し、その信頼性を損なう。
以前の作業は、LVLMが言語事前に過度に依存していることと、ロジットキャリブレーションによってそれを緩和しようとすることによる。
我々は,LVLMがオブジェクトの存在の信頼性を忠実に検証できるように,Language-Prior-Free Verificationを提案する。
論文 参考訳(メタデータ) (2026-01-30T01:37:53Z) - Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis [0.42970700836450487]
我々は、幻覚は最適化の失敗ではなく、トランスフォーマーモデルのアーキテクチャ上の必然性であると主張している。
実験の結果,幻覚は,外的真理検証と禁忌モジュールによってのみ除去できることが示唆された。
幻覚は生成的アーキテクチャの構造的特性であると結論付けている。
論文 参考訳(メタデータ) (2025-12-16T17:39:45Z) - Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models [12.747507415841168]
本研究では,制約付き知識領域における幻覚の因果関係について,チェーン・オブ・ソート(Chain-of-Thought)の軌跡を監査することによって検討する。
我々の分析によると、長いCoT設定では、RLLMは欠陥のある反射的推論を通じてバイアスやエラーを反復的に補強することができる。
驚いたことに、幻覚の原因の直接的な介入でさえも、連鎖が「連鎖不規則性」を示すため、その効果を覆すことができないことが多い。
論文 参考訳(メタデータ) (2025-05-19T14:11:09Z) - HalluLens: LLM Hallucination Benchmark [49.170128733508335]
大規模言語モデル(LLM)は、しばしばユーザ入力やトレーニングデータから逸脱する応答を生成する。
本稿では,新たな内因性評価タスクと既存内因性評価タスクを併用した総合幻覚ベンチマークを提案する。
論文 参考訳(メタデータ) (2025-04-24T13:40:27Z) - Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling [78.78822033285938]
VLM(Vision-Language Models)は視覚的理解に優れ、視覚幻覚に悩まされることが多い。
本研究では,幻覚を意識したトレーニングとオンザフライの自己検証を統合した統合フレームワークREVERSEを紹介する。
論文 参考訳(メタデータ) (2025-04-17T17:59:22Z) - The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination [85.18584652829799]
本稿では,知識のシェードイングをモデル化することで,事実の幻覚を定量化する新しい枠組みを提案する。
オーバシャドウ(27.9%)、MemoTrap(13.1%)、NQ-Swap(18.3%)のモデル事実性を顕著に向上させる。
論文 参考訳(メタデータ) (2025-02-22T08:36:06Z) - Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer [51.7407540261676]
本研究では,モデルが常に正しい解答を行うことのできる幻覚の別のタイプについて検討するが,一見自明な摂動は,高い確実性で幻覚応答を生じさせる。
この現象は特に医学や法学などの高度な領域において、モデルの確実性はしばしば信頼性の代用として使用される。
CHOKEの例は、プロンプト間で一貫性があり、異なるモデルやデータセットで発生し、他の幻覚と根本的に異なることを示す。
論文 参考訳(メタデータ) (2025-02-18T15:46:31Z) - Don't Say What You Don't Know: Improving the Consistency of Abstractive
Summarization by Constraining Beam Search [54.286450484332505]
本研究は,幻覚とトレーニングデータの関連性を解析し,学習対象の要約を学習した結果,モデルが幻覚を呈する証拠を見出した。
本稿では,ビーム探索を制約して幻覚を回避し,変換器をベースとした抽象要約器の整合性を向上させる新しい復号法であるPINOCCHIOを提案する。
論文 参考訳(メタデータ) (2022-03-16T07:13:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。