論文の概要: Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
- arxiv url: http://arxiv.org/abs/2603.25555v1
- Date: Thu, 26 Mar 2026 15:27:27 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-27 20:52:48.358061
- Title: Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
- Title(参考訳): マルチモーダル画像融合による眼科手術における総合的リアルタイムシーン理解に向けて
- Authors: Nikolo Rohrmoser, Ghazal Ghazaei, Michael Sommersperger, Nassir Navab,
- Abstract要約: 本稿では,共同機器検出,キーポイントの局所化,ツール間距離推定を行うための時間的,リアルタイムなネットワークアーキテクチャを提案する。
実験では、信頼性の高い機器のローカライゼーションとキーポイント検出(95.79% mAP50)が示され、i OCTの組み入れにより、ツール・タスク距離の推定が大幅に改善された。
- 参考スコア(独自算出の注目度): 40.97183044655793
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Purpose: The integration of multimodal imaging into operating rooms paves the way for comprehensive surgical scene understanding. In ophthalmic surgery, by now, two complementary imaging modalities are available: operating microscope (OPMI) imaging and real-time intraoperative optical coherence tomography (iOCT). This first work toward temporal OPMI and iOCT feature fusion demonstrates the potential of multimodal image processing for multi-head prediction through the example of precise instrument tracking in vitreoretinal surgery. Methods: We propose a multimodal, temporal, real-time capable network architecture to perform joint instrument detection, keypoint localization, and tool-tissue distance estimation. Our network design integrates a cross-attention fusion module to merge OPMI and iOCT image features, which are efficiently extracted via a YoloNAS and a CNN encoder, respectively. Furthermore, a region-based recurrent module leverages temporal coherence. Results: Our experiments demonstrate reliable instrument localization and keypoint detection (95.79% mAP50) and show that the incorporation of iOCT significantly improves tool-tissue distance estimation, while achieving real-time processing rates of 22.5 ms per frame. Especially for close distances to the retina (below 1 mm), the distance estimation accuracy improved from 284 $μm$ (OPMI only) to 33 $μm$ (multimodal). Conclusion: Feature fusion of multimodal imaging can enhance multi-task prediction accuracy compared to single-modality processing and real-time processing performance can be achieved through tailored network design. While our results demonstrate the potential of multi-modal processing for image-guided vitreoretinal surgery, they also underline key challenges that motivate future research toward more reliable, consistent, and comprehensive surgical scene understanding.
- Abstract(参考訳): 目的: 手術室へのマルチモーダルイメージングの統合は, 総合的な手術シーン理解の道を開く。
現在、眼科手術では、手術顕微鏡(OPMI)とリアルタイム術中光コヒーレンス断層撮影(iOCT)の2つの相補的な画像モダリティが利用可能である。
本研究は, 硝子体手術における計測器追跡の例を例として, 時間的OPMIとiOCT機能融合に向けた最初の研究により, マルチヘッド予測のためのマルチモーダル画像処理の可能性を示すものである。
方法: 複数モーダル, 時間, リアルタイムなネットワークアーキテクチャを提案し, 共同計測, キーポイントの局所化, ツール・タスク距離推定を行う。
ネットワーク設計では,OPMI と iOCT の画像を融合するための相互注意融合モジュールを統合し,それぞれが YoloNAS と CNN エンコーダによって効率的に抽出される。
さらに、リージョンベースのリカレントモジュールは時間的コヒーレンスを利用する。
結果: 実験では, 信頼性の高い機器の局所化とキーポイント検出(95.79% mAP50)を実証し, iOCTの組み込みにより, 1フレームあたりのリアルタイム処理速度が22.5msとなるとともに, 工具間距離推定が大幅に向上することを示した。
特に網膜との近接距離(1mm以下)では、距離推定精度は284$μm$ (OPMIのみ) から33$μm$ (multimodal) に改善された。
結論:マルチモーダル画像の特徴融合により,単一モーダル処理と比較してマルチタスク予測精度が向上し,ネットワーク設計の調整によりリアルタイム処理性能が向上する。
画像ガイド下硝子体手術におけるマルチモーダルプロセッシングの可能性を示すとともに,より信頼性,一貫性,総合的な手術シーン理解に向けた今後の研究を動機づける重要な課題も明らかにした。
関連論文リスト
- Unsupervised MRI-US Multimodal Image Registration with Multilevel Correlation Pyramidal Optimization [14.509109797489499]
マルチレベル相関ピラミッド最適化(MCPO)に基づく教師なしマルチモーダル医用画像登録手法を提案する。
提案手法は,ReMIND2Regの検証フェーズとテストフェーズにおける第1位を達成する。
本手法が術前から術中画像登録に広く適用可能であることを示す。
論文 参考訳(メタデータ) (2026-02-06T01:03:57Z) - A Unified Model for Compressed Sensing MRI Across Undersampling Patterns [69.19631302047569]
様々な計測アンサンプパターンと画像解像度に頑健な統合MRI再構成モデルを提案する。
我々のモデルは、拡散法よりも600$times$高速な推論で、最先端CNN(End-to-End VarNet)の4dBでSSIMを11%改善し、PSNRを4dB改善する。
論文 参考訳(メタデータ) (2024-10-05T20:03:57Z) - Friends Across Time: Multi-Scale Action Segmentation Transformer for
Surgical Phase Recognition [2.10407185597278]
オフライン手術相認識のためのMS-AST(Multi-Scale Action Causal Transformer)とオンライン手術相認識のためのMS-ASCT(Multi-Scale Action Causal Transformer)を提案する。
オンラインおよびオフラインの外科的位相認識のためのColec80データセットでは,95.26%,96.15%の精度が得られる。
論文 参考訳(メタデータ) (2024-01-22T01:34:03Z) - Phase-Specific Augmented Reality Guidance for Microscopic Cataract
Surgery Using Long-Short Spatiotemporal Aggregation Transformer [14.568834378003707]
乳化白内障手術(英: Phaemulsification cataract surgery, PCS)は、外科顕微鏡を用いた外科手術である。
PCS誘導システムは、手術用顕微鏡映像から貴重な情報を抽出し、熟練度を高める。
既存のPCSガイダンスシステムでは、位相特異なガイダンスに悩まされ、冗長な視覚情報に繋がる。
本稿では,認識された手術段階に対応するAR情報を提供する,新しい位相特異的拡張現実(AR)誘導システムを提案する。
論文 参考訳(メタデータ) (2023-09-11T02:56:56Z) - LoViT: Long Video Transformer for Surgical Phase Recognition [59.06812739441785]
短時間・長期の時間情報を融合する2段階のLong Video Transformer(LoViT)を提案する。
このアプローチは、Colec80とAutoLaparoデータセットの最先端メソッドを一貫して上回る。
論文 参考訳(メタデータ) (2023-05-15T20:06:14Z) - Multi-modal Aggregation Network for Fast MR Imaging [85.25000133194762]
我々は,完全サンプル化された補助モダリティから補完表現を発見できる,MANetという新しいマルチモーダル・アグリゲーション・ネットワークを提案する。
我々のMANetでは,完全サンプリングされた補助的およびアンアンサンプされた目標モダリティの表現は,特定のネットワークを介して独立に学習される。
私たちのMANetは、$k$-spaceドメインの周波数信号を同時に回復できるハイブリッドドメイン学習フレームワークに従います。
論文 参考訳(メタデータ) (2021-10-15T13:16:59Z) - Patch-based field-of-view matching in multi-modal images for
electroporation-based ablations [0.6285581681015912]
マルチモーダルイメージングセンサーは、現在、介入治療作業フローの異なるステップに関与している。
この情報を統合するには、取得した画像間の観測された解剖の正確な空間的アライメントに依存する。
本稿では, ボクセルパッチを用いた地域登録手法が, ボクセルワイドアプローチと「グローバルシフト」アプローチとの間に優れた構造的妥協をもたらすことを示す。
論文 参考訳(メタデータ) (2020-11-09T11:27:45Z) - Multifold Acceleration of Diffusion MRI via Slice-Interleaved Diffusion
Encoding (SIDE) [50.65891535040752]
本稿では,Slice-Interleaved Diffusionと呼ばれる拡散符号化方式を提案する。
SIDEは、拡散重み付き(DW)画像ボリュームを異なる拡散勾配で符号化したスライスでインターリーブする。
また,高いスライスアンサンプデータからDW画像を効果的に再構成するためのディープラーニングに基づく手法を提案する。
論文 参考訳(メタデータ) (2020-02-25T14:48:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。