論文の概要: CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
- arxiv url: http://arxiv.org/abs/2606.27264v1
- Date: Thu, 25 Jun 2026 16:47:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-26 18:46:32.361042
- Title: CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
- Title(参考訳): CORTEX: 信頼できる3D胸部CTMLLMのための構造化推論ベンチマーク
- Authors: Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, Christoph Lippert,
- Abstract要約: 我々は3次元胸部CTのための構造化推論ベンチマークであるCORTEXを紹介する。
各質問に対して、CORTEXは欠落した推論を4段階の診断トレースとして復元する。
CORTEXは、オープンエンドVQA、クローズドエンドVQA、レポート生成を含む76,177個の検証された推論トレースからなる。
- 参考スコア(独自算出の注目度): 35.41351101286175
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer, making it hard to interpret and verify, especially in 3D radiology, where a diagnosis should be traceable to evidence in the scan. Existing chest CT question-answering datasets compound this by reducing expert radiology reports to answer-only pairs, dropping the reasoning that links findings to conclusions and omitting the patient history clinicians rely on. As a result, reasoning-capable 3D chest CT MLLMs remain out of reach, as neither the structured supervision needed to train them nor the protocol needed to verify their reasoning yet exists. We introduce CORTEX (Clinically Organized Reasoning and sTructured EXplanation), a structured reasoning benchmark for 3D chest CT. For each question, CORTEX restores the missing reasoning as a four-stage diagnostic trace mirroring a radiologist's workflow: task understanding, visual observation, diagnostic reasoning, and answer synthesis. We generate these traces using frontier large language models with broad medical and general-domain knowledge, then filter and verify them with a stage-level evaluation protocol combining automated rubric scoring with expert radiologist review. Crucially, both the reasoning structure and evaluation rubrics are designed in close collaboration with clinicians. Built on CT-RATE, a large, publicly available chest CT dataset without reasoning annotations, CORTEX comprises 76,177 validated reasoning traces across open-ended VQA, closed-ended VQA, and report generation, providing both the structured supervision and the stage-level evaluation protocol needed to build and evaluate trustworthy reasoning models for 3D chest CT. Our dataset and evaluation code will be made publicly available upon acceptance.
- Abstract(参考訳): マルチモーダル大言語モデル (MLLM) における推論は, 医用画像における強い将来性を示している。
しかしながら、この推論は通常、最終回答によってのみ判断される自由形式のテキストであり、特に3Dラジオグラフィーでは、スキャン中の証拠に診断が追跡できるため、解釈と検証が困難である。
既存の胸部CT問合せデータセットは、専門家の放射線医学レポートを回答のみのペアに減らし、結果と結論を関連付ける理由を排除し、患者の歴史臨床医が頼りにする。
その結果、推論可能な3D胸部CT MLLMは、トレーニングに必要な構造的な監督や、推論の検証に必要なプロトコルが存在しないため、到達できないままとなった。
我々は3次元胸部CTのための構造化推論ベンチマークであるCORTEX(Clinically Organized Reasoning and sTructured Explanation)を紹介する。
それぞれの質問に対して、CORTEXは4段階の診断トレースとして、タスク理解、視覚観察、診断推論、回答合成という、放射線学者のワークフローを反映した行方不明の推論を復元する。
広い医学知識と一般ドメイン知識を持つフロンティア大言語モデルを用いてこれらのトレースを生成し、その後、専門家の放射線学者のレビューと自動ルーリックスコアを組み合わせたステージレベルの評価プロトコルを用いてフィルタリングし、検証する。
重要なことは、推論構造と評価ルーブリックの両方が、臨床医との密接なコラボレーションのために設計されている。
CORTEXはCT-RATEをベースとしており、3D胸部CTの信頼できる推論モデルの構築と評価に必要な構造化された監視とステージレベルの評価プロトコルの両方を提供する、76,177個の検証済みのVQA、クローズドエンドVQA、レポート生成を含んでいる。
私たちのデータセットと評価コードは、受理時に公開されます。
関連論文リスト
- Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging [10.149461900482192]
医用画像上での視覚言語モデル (VLM) の評価には、臨床的に基礎を置き、拡張性があり、評価のために制御されるベンチマークが必要である。
本稿では,2つのプライベートラジオグラフィーレポートと3Dオンコロジーイメージングから直接,複数選択VQAデータセットを生成する自動エージェント駆動パイプラインを提案する。
論文 参考訳(メタデータ) (2026-06-01T19:27:42Z) - Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models [3.6659638651482536]
近年の3次元医用視覚言語モデルの進歩により、ボリューム画像やテキストに対する共同推論が可能になった。
これらのモデルが3次元ボリュームから空間的基底解剖学を学習したのか、あるいは主に学習前と言語相関に頼っているのかは、まだ不明である。
我々は3次元CTデータにおける意味空間的推論を評価するためのベンチマークであるCT-SpatialVQAを紹介する。
論文 参考訳(メタデータ) (2026-05-09T08:16:00Z) - MedScribe: Clinically Grounded CT Reporting through Agentic Workflows [13.40306812882295]
視覚言語モデル(VLM)は、自動放射線診断レポート生成の可能性を示している。
我々は,仮説駆動型フレームワークであるMedScribeを紹介し,レポート生成を反復的証拠取得プロセスとして再構築する。
論文 参考訳(メタデータ) (2026-05-03T08:32:40Z) - EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT [29.0378459959757]
EXACTは3次元胸部CTの異常認識基盤モデルである。
2つの臨床スキャンと放射線学レポートから空間的に解決された表現を学習する。
EXACTは臨床的に関係のあるCTタスクに対して一貫した改善を示す。
論文 参考訳(メタデータ) (2026-04-27T07:57:47Z) - CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation [51.11942945171396]
従来の評価指標は、語彙重なり合いやエンティティマッチングの粗い尺度のみを提供する。
我々はCT-RATEとMerlinのベンチマークであるCT-FineBenchを提案し、CTレポートの微細な事実整合性を評価する。
我々のベンチマークは、綿密な質問回答(QA)ベースのプロセスによって構築されます。
論文 参考訳(メタデータ) (2026-04-27T03:32:46Z) - M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding [66.78251988482222]
CoT(Chain-of-Thought)推論は、ステップバイステップの中間推論を奨励することによって、大規模言語モデルの強化に有効であることが証明されている。
医用画像理解のための現在のベンチマークでは、推論パスを無視しながら最終回答に重点を置いている。
M3CoTBenchは、透明で信頼性が高く、診断的に正確な医療用AIシステムの開発を促進することを目的としている。
論文 参考訳(メタデータ) (2026-01-13T17:42:27Z) - MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement [1.6355783973385114]
多視点認識知識強化型TansfoRmer(MvKeTR)
複数の解剖学的ビューから診断情報を効果的に合成するために、ビューアウェアのMVPAを提案する。
クエリボリュームに基づいて、最も類似したレポートを取得するために、Cross-Modal Knowledge Enhancer (CMKE) が考案されている。
論文 参考訳(メタデータ) (2024-11-27T12:58:23Z) - 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models [51.855377054763345]
本稿では,VQAに基づく医用視覚言語モデルである3D-CT-GPTについて紹介する。
パブリックデータセットとプライベートデータセットの両方の実験により、3D-CT-GPTはレポートの正確さと品質という点で既存の手法を著しく上回っていることが示された。
論文 参考訳(メタデータ) (2024-09-28T12:31:07Z) - RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis [56.57177181778517]
RadGenome-Chest CTはCT-RATEに基づく大規模3次元胸部CT解釈データセットである。
私たちは、最新の強力なユニバーサルセグメンテーションと大きな言語モデルを活用して、元のデータセットを拡張します。
論文 参考訳(メタデータ) (2024-04-25T17:11:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。