論文の概要: Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
- arxiv url: http://arxiv.org/abs/2609.00866v1
- Date: Tue, 01 Sep 2026 08:01:31 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.44485
- Title: Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
- Title(参考訳): 自動病理診断のためのベンチマークビジョンランゲージモデルとレポート生成
- Authors: Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balcı, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan Sökmen, Ece Tuğba Cebeci, Ahmet Halıcı, Musa Balcı, Kardelen Peçenek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn,
- Abstract要約: 本報告では,5施設から約10,500対の医療用パンアジアWSIを報告した。
マルチモーダルモデルの体系的評価のためのMICCAIチャレンジを通じてREG 2025ベンチマークを確立する。
- 参考スコア(独自算出の注目度): 28.553483401938195
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.
- Abstract(参考訳): 視覚言語モデル(VLM)の急速な進歩は、計算病理学の進歩を加速させているが、大規模なWSIレポートデータセットの不足や、空間的に分散した視覚パターンを構造化された臨床テキストにマッピングする複雑さによって、全スライディング画像(WSI)ベースの病態レポートの生成は制限されている。
そこで本研究では,5施設から約10,500対のPan-Asia WSI-Reportデータセットを導入し,マルチモーダルモデルの体系的評価のためのMICCAIチャレンジを通じてREG 2025ベンチマークを構築した。
我々は,事前学習されたVLM,マルチインスタンス学習フレームワーク,階層的エキスパートモデル,検索拡張生成,モーダル変換を対象とする提案手法を解析した。
以上の結果から, VLM単独の使用が優れた性能に十分であることを示すのではなく, 構造化された報告表現, 階層的診断分解, 効果的なマルチモーダルグラウンドディングの恩恵を受けることが示唆された。
定量的属性推定の不安定性 (例えば, 数値幻覚) や, 診断過小評価の傾向など, いくつかの誤りは診断の落とし穴に類似している。
これらの結果は,WSI に基づく構造化レポート生成とコンピュータ病理学における視覚言語理解を評価するためのベンチマークとして REG 2025 を確立し,臨床基盤のマルチモーダル病理学モデルの設計に関する知見を提供する。
関連論文リスト
- Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology [54.42160154368719]
WSI(Whole-Slide Image)は、病理のデジタルトランスフォーメーションを可能にする。
病理組織像と病理所見を共同で解析するマルチモーダルモデルは、自動病理報告生成とAIによる診断に有望な可能性を示唆している。
本稿では,WSIsを併用した子宮疾患データセットであるTUM-Uteriaを紹介する。
論文 参考訳(メタデータ) (2026-07-04T20:34:59Z) - Multimodal Graph-based Classification of Esophageal Motility Disorders [73.90451172929117]
食道運動障害の診断は,高分解能インピーダンス測定データの複雑化と臨床解釈の多様性が原因で大きな課題となる。
本研究は,HRIM記録と患者固有の情報を組み合わせたマルチモーダル機械学習に基づく分類手法の実現可能性について検討し,食道生理学のグラフベースモデリングを取り入れた。
論文 参考訳(メタデータ) (2026-05-13T14:52:12Z) - Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs [63.535652574541764]
MLLM(Multimodal Large Language Models)は医用画像解析において顕著な可能性を示した。
消化器内視鏡におけるそれらの応用は、現在、2つの重要な限界によって妨げられている。
本稿では,これらの課題に対処する新しい臨床認知アライメント(CogAlign)フレームワークを提案する。
論文 参考訳(メタデータ) (2026-03-21T07:47:37Z) - A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis [82.01597026329158]
本稿では,組織合成のための相関調整フレームワーク(CRAFTS)について紹介する。
CRAFTSは、生物学的精度を確保するためにセマンティックドリフトを抑制する新しいアライメント機構を組み込んでいる。
本モデルは,30種類の癌にまたがる多彩な病理像を生成する。
論文 参考訳(メタデータ) (2025-12-15T10:22:43Z) - PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology [3.459714932882085]
現在の視覚言語(VL)モデルは、構造化された病理報告の解釈に必要な複雑な推論を捉えるのに苦労することが多い。
病理領域内での階層的意味理解と構成的推論におけるVLモデルの能力を評価するために設計された新しいベンチマークであるPathoHR-Benchを提案する。
さらに、マルチモーダルコントラスト学習のための拡張および摂動サンプルを生成する、病理特異的なVLトレーニングスキームを導入する。
論文 参考訳(メタデータ) (2025-09-07T15:42:38Z) - AMRG: Extend Vision Language Models for Automatic Mammography Report Generation [4.366802575084445]
マンモグラフィーレポート生成は、医療AIにおいて重要で未発見の課題である。
マンモグラフィーレポートを生成するための最初のエンドツーエンドフレームワークであるAMRGを紹介する。
DMIDを用いた高分解能マンモグラフィーと診断レポートの公開データセットであるAMRGのトレーニングと評価を行った。
論文 参考訳(メタデータ) (2025-08-12T06:37:41Z) - Clinical-grade Multi-Organ Pathology Report Generation for Multi-scale Whole Slide Images via a Semantically Guided Medical Text Foundation Model [3.356716093747221]
患者に対する病理報告を生成するために, 患者レベル多臓器報告生成(PMPRG)モデルを提案する。
我々のモデルはMETEORスコア0.68を達成し、我々のアプローチの有効性を実証した。
論文 参考訳(メタデータ) (2024-09-23T22:22:32Z) - PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology [7.87900104748629]
6つの異なるタスクをカバーする約45,000のケースのデータセットを慎重にコンパイルしました。
特にLLaVA, Qwen-VL, InternLMを微調整したマルチモーダル大規模モデルで, このデータセットを用いて命令ベースの性能を向上させる。
論文 参考訳(メタデータ) (2024-08-13T17:05:06Z) - Potential of Multimodal Large Language Models for Data Mining of Medical Images and Free-text Reports [51.45762396192655]
特にGemini-Vision-Series (Gemini) と GPT-4-Series (GPT-4) は、コンピュータビジョンのための人工知能のパラダイムシフトを象徴している。
本研究は,14の医用画像データセットを対象に,Gemini,GPT-4,および4つの一般的な大規模モデルの性能評価を行った。
論文 参考訳(メタデータ) (2024-07-08T09:08:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。