論文の概要: NV-Reason-CT: 3D Visual Language Model for CT Analysis
- arxiv url: http://arxiv.org/abs/2609.27511v2
- Date: Fri, 25 Sep 2026 01:24:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-28 18:27:51.041448
- Title: NV-Reason-CT: 3D Visual Language Model for CT Analysis
- Title(参考訳): NV-Reason-CT:CT解析のための3次元視覚言語モデル
- Abstract要約: NV-Reason-CTは胸部CTおよび腹部CTのための生成的視覚言語モデルである。
このモデルは、ネイティブな3Dビジョントランスフォーマーと言語モデルとを結合し、すべての視覚トークンとその明示的な3D座標を言語デコーディングに渡す。
我々は70,111個のCT画像入力から約550,000個のマルチモーダル・インストラクション例をキュレートしたコーパスで訓練する。
- 参考スコア(独自算出の注目度): 21.378661110452725
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.
- Abstract(参考訳): NV-Reason-CT, 胸部・腹部CT生成モデル, ネイティブ3次元ビジュアルエンコーディングと放射線医誘導推論を併用した。
このモデルは、ネイティブな3D視覚変換器と言語モデルとを結合し、全ての視覚トークンとその明示的な3D座標を、さらなる空間トークンのマージなしに言語デコーディングに渡す。
これは、視覚エンコーダ内の体積空間情報と、テキストとの共同処理中に言語モデルの位置エンコーディングを通して保持する。
我々は70,111個のCT画像入力から約550,000個のマルチモーダル・インストラクション・サンプルを収集し、標準化されたレポート、異常や解剖学的特異な質問、マルチターン・インタラクション、および記録および転写された専門的CT解釈からの放射線技師による推論を組み合わせる。
エキスパートアノテーションは、直接の監視と、追加のレポート地上合成推論のガイドを提供する。
終末監視細調整 (SFT) に続いて, グループ相対政策最適化 (GRPO) が施行され, 胸部, 腹部の異常セットに対する報酬が検証された。
このモデルは、異常分類、レポート生成、レビュー可能な観察、微分診断、不確実性を含む対話的推論をサポートする。
評価は、公開CTベンチマークと保留のNIHコホートにまたがる。
CT-RATEでは、NV-Reason-CTはタスク固有の分類ヘッドなしで0.614のマクロF1と0.871のマクロAUROCを達成し、生成されたレポートは0.592のレポート由来のマクロF1を達成する。
専門家の放射線学者による予備的な研究で、AI支援レビューは良好な信頼評価を受け、平均報告された解釈と報告時間の50%削減に結びついた。
本稿では,医療画像のための説明可能なAIに関する再現可能な研究を支援するためのモデルとトレーニングコードをリリースする。
関連論文リスト
- Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography [75.7789001619107]
ビジョン言語による事前トレーニングは、汎用医療AIにとって大きな可能性を秘めている。
3つの主要なコンポーネントを特徴とするフレームワークを提案する。
診断対応プロンプト戦略では、トレーニング前の推論ギャップを埋めるために、実際の臨床用語と集約された疾患プロトタイプを用いている。
論文 参考訳(メタデータ) (2026-06-24T08:24:45Z) - EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT [29.0378459959757]
EXACTは3次元胸部CTの異常認識基盤モデルである。
2つの臨床スキャンと放射線学レポートから空間的に解決された表現を学習する。
EXACTは臨床的に関係のあるCTタスクに対して一貫した改善を示す。
論文 参考訳(メタデータ) (2026-04-27T07:57:47Z) - Towards a Holistic Framework for Multimodal Large Language Models in Three-dimensional Brain CT Report Generation [42.06416052431378]
2Dラジオグラフィーキャプションは、ボリューム3D解剖学における現実の診断課題を反映するものではない。
我々は18,885組の3D-BrainCTデータセットを収集し,臨床ビジュアルインストラクション・チューニングを用いて,脳波モデルを用いて放射線治療を施した3D脳CTレポートを作成した。
私たちの研究は、3Dの脳CTデータセットのキュレーション、微調整による解剖学的意味のある言語モデル、堅牢な放射線学評価指標の提案など、総合的な枠組みを具現化したものです。
論文 参考訳(メタデータ) (2024-07-02T12:58:35Z) - RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis [56.57177181778517]
RadGenome-Chest CTはCT-RATEに基づく大規模3次元胸部CT解釈データセットである。
私たちは、最新の強力なユニバーサルセグメンテーションと大きな言語モデルを活用して、元のデータセットを拡張します。
論文 参考訳(メタデータ) (2024-04-25T17:11:37Z) - CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios [53.94122089629544]
我々は,CT-GLIP(Grounded Language- Image Pretraining with CT scans)を導入する。
本手法は,104臓器にわたる17,702症例を対象に,44,011例の臓器レベルの視覚テキストペアからなるマルチモーダルCTデータセットを用いて訓練し,自然言語を用いて臓器と異常をゼロショットで識別できることを実証した。
論文 参考訳(メタデータ) (2024-04-23T17:59:01Z) - Auxiliary Signal-Guided Knowledge Encoder-Decoder for Medical Report
Generation [107.3538598876467]
放射線技師の動作パターンを模倣する補助信号誘導知識デコーダ(ASGK)を提案する。
ASGKは、内的特徴融合と外部医療言語情報を統合して、医療知識の伝達と学習をガイドする。
論文 参考訳(メタデータ) (2020-06-06T01:00:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。