論文の概要: Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints
- arxiv url: http://arxiv.org/abs/2606.13211v1
- Date: Thu, 11 Jun 2026 11:19:31 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-12 15:55:27.754589
- Title: Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints
- Title(参考訳): 医用画像AIにおける幻覚 : 規制下での分類・検出・緩和のためのクロスモーダル分析フレームワーク
- Authors: Omar Alshahrani, Muzammil Behzad,
- Abstract要約: 幻覚は臨床的にはあり得ないが、実際には正しくない出力である。
汎用基礎モデルは幻覚特異的ベンチマークにおいて医療特化モデルより優れている。
FDAのライフサイクル監視は依然として不可欠である。
- 参考スコア(独自算出の注目度): 0.3437656066916039
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: AI systems are being deployed across medical imaging faster than their failure modes are understood. At this point in time, the failure of greatest clinical concern is hallucination: clinically plausible but factually incorrect outputs, including fabricated anatomical structures, missed findings, incorrect laterality, and invented measurements in generated reports, with direct consequences, for example, for biopsy decisions, staging, and treatment planning. This structured narrative synthesizes peer-reviewed studies, benchmark datasets, and FDA regulatory guidance across five imaging modalities to produce a cross-modality analysis of hallucination taxonomy, etiology, detection, and mitigation. Specifically, we address three questions in this study: (1) how can existing taxonomies be unified across modalities?, (2) how do medical-specialized foundation models hallucinate less than general-purpose ones?, and (3) which mitigation strategies are effective and compatible with FDA lifecycle oversight? We note that three taxonomic frameworks together cover the imaging pipeline in a way no single framework does alone. We also highlight that general-purpose foundation models outperform medical-specialized models on hallucination-specific benchmarks, indicating that narrow domain fine-tuning can introduce overfitting-induced confabulation. At the same time, the oversight of radiologists remains essential; for instance, a very high percentage of of AI-generated flags required expert correction before clinical use. Physics-informed architectural constraints, Chain-of-Thought prompting, and human-in-the-loop safeguards each address different failure modes and is effective when combined. All findings are mapped to the FDA's Total Product Lifecycle and Predetermined Change Control Plan frameworks, which treat hallucination management as a lifecycle obligation rather than a pre-deployment checklist.
- Abstract(参考訳): AIシステムは、その障害モードが理解されるよりも早く、医療画像にデプロイされている。
この時点で、最大の臨床的関心事の失敗は幻覚である: 臨床的に妥当だが事実に誤りのあるアウトプット: 製造された解剖学的構造、発見の欠如、不正確な横行性、および生成されたレポートで発明された測定結果、例えば生検決定、ステージング、治療計画などである。
この構造化された物語は、5つの画像モダリティにわたるピアレビューされた研究、ベンチマークデータセット、FDAの規制ガイダンスを合成し、幻覚分類学、エチオロジー、検出、緩和の相互モダリティ分析を生成する。
具体的には,(1)既存の分類体系をモダリティ全体にわたって一体化するにはどうすればよいのか,という3つの疑問に対処する。
, (2)医学専門化基礎モデルでは, 一般のモデルよりも幻覚が低いか?
および(3) FDAのライフサイクル監視に有効で相容れない緩和策は何か。
3つの分類学的フレームワークが一緒に、単一のフレームワークだけではできない方法で、イメージングパイプラインをカバーすることに留意する。
また, 幻覚特異的ベンチマークにおいて, 汎用基礎モデルは医療特化モデルよりも優れており, 狭い領域の微調整が過度に適合する折り畳みを生じさせる可能性が示唆された。
同時に、放射線技師の監督は依然として不可欠であり、例えば、AIが生成するフラグの非常に高い割合は、臨床使用の前に専門家の修正を必要とした。
物理インフォームドアーキテクチャ制約、Chain-of-Thoughtプロンプト、Human-in-the-loopセーフガードは、それぞれ異なる障害モードに対処し、組み合わせると有効である。
すべての発見は、FDAのトータル・プロダクト・ライフサイクルと事前決定された変更制御計画のフレームワークにマッピングされ、幻覚管理を事前デプロイチェックリストではなくライフサイクル義務として扱う。
関連論文リスト
- Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs [53.697403481143404]
LVLM(Large Vision-Language Models)は、医用画像処理タスクにおいて強力なパフォーマンスを実現している。
しかし、実際の不整合、視力の低下、臨床的に有意義なフィードバックによる不一致が続く傾向にある。
論文 参考訳(メタデータ) (2026-06-10T18:35:36Z) - MedScribe: Clinically Grounded CT Reporting through Agentic Workflows [13.40306812882295]
視覚言語モデル(VLM)は、自動放射線診断レポート生成の可能性を示している。
我々は,仮説駆動型フレームワークであるMedScribeを紹介し,レポート生成を反復的証拠取得プロセスとして再構築する。
論文 参考訳(メタデータ) (2026-05-03T08:32:40Z) - Bridging Restoration and Diagnosis: A Comprehensive Benchmark for Retinal Fundus Enhancement [13.241000354696903]
EyeBench-V2は、拡張モデルの性能と臨床的有用性の間のギャップを埋めるために設計されたベンチマークである。
我々のベンチマークは、既存の生成モデルの厳密なタスク指向の分析を提供する。
論文 参考訳(メタデータ) (2026-04-04T17:24:23Z) - Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models [62.932580559941414]
VLM(Vision-Language Models)は、しばしば「ハロシン化(hallucinate)」する。
本稿では,静的な出力誤差からモデル計算認知の動的病理へ再キャストし,幻覚を診断するための新しいパラダイムを提案する。
論文 参考訳(メタデータ) (2026-03-16T17:20:38Z) - PathReasoner-R1: Instilling Structured Reasoning into Pathology Vision-Language Model via Knowledge-Guided Policy Optimization [6.821738567680833]
PathReasonerは,WSI推論の最初の大規模データセットである。
PathReasoner-R1は、教師付き微調整と推論指向の強化学習を相乗し、構造化されたチェーン・オブ・シント機能を注入する。
実験により、PathReasoner-R1はPathReasonerと公開ベンチマークの両方で、様々な画像スケールで最先端のパフォーマンスを達成することが示された。
論文 参考訳(メタデータ) (2026-01-29T12:21:16Z) - Aligning Findings with Diagnosis: A Self-Consistent Reinforcement Learning Framework for Trustworthy Radiology Reporting [37.57009831483529]
MLLM(Multimodal Large Language Models)は放射線学レポート生成に強い可能性を示している。
本フレームワークは, より詳細な発見のための思考ブロックと, 構造化された疾患ラベルに対する回答ブロックという, 生成を2つの異なる構成要素に再構成する。
論文 参考訳(メタデータ) (2026-01-06T14:17:44Z) - A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis [82.01597026329158]
本稿では,組織合成のための相関調整フレームワーク(CRAFTS)について紹介する。
CRAFTSは、生物学的精度を確保するためにセマンティックドリフトを抑制する新しいアライメント機構を組み込んでいる。
本モデルは,30種類の癌にまたがる多彩な病理像を生成する。
論文 参考訳(メタデータ) (2025-12-15T10:22:43Z) - Medical Hallucinations in Foundation Models and Their Impact on Healthcare [71.15392179084428]
基礎モデルの幻覚は自己回帰訓練の目的から生じる。
トップパフォーマンスモデルは、チェーン・オブ・シークレット・プロンプトで強化された場合、97%の精度を達成した。
論文 参考訳(メタデータ) (2025-02-26T02:30:44Z) - Cross-modal Clinical Graph Transformer for Ophthalmic Report Generation [116.87918100031153]
眼科報告生成(ORG)のためのクロスモーダルな臨床グラフ変換器(CGT)を提案する。
CGTは、デコード手順を駆動する事前知識として、臨床関係を視覚特徴に注入する。
大規模FFA-IRベンチマークの実験は、提案したCGTが従来のベンチマーク手法より優れていることを示した。
論文 参考訳(メタデータ) (2022-06-04T13:16:30Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。