論文の概要: Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery
- arxiv url: http://arxiv.org/abs/2608.31065v1
- Date: Mon, 31 Aug 2026 16:41:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-01 18:31:31.519018
- Title: Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery
- Title(参考訳): 硝子体手術における相認識のためのナレーション,顕微鏡およびiOCT画像の多モード共有潜在表現
- Abstract要約: 硝子体手術における文脈認識型コンピュータによるフィードバックには,外科的位相認識が重要である。
手術用ナレーションと術中OCTを橋渡しするための共有アンカーとして顕微鏡ビューを用いたフレームワークを提案する。
これは、顕微鏡視、i OCT Bスキャン、および外科的位相認識のための共有潜在空間における外科的ナレーションを統一する最初の試みである。
- 参考スコア(独自算出の注目度): 37.94518682851121
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Surgical phase recognition is key to context-aware computer-assisted feedback in vitreoretinal procedures, yet the scarcity of synchronized multimodal intraoperative data, particularly microscope views and intraoperative OCT, limits approaches that aim to replicate the multimodal integration surgeons perform naturally. Surgical narration, by contrast, is abundantly available online and offers rich semantic supervision. Prior work has mainly explored pairwise contrastive learning (e.g., intraoperative OCT-microscope or microscope-narration), leaving the joint modeling of all three modalities largely unexplored. We introduce a framework that uses microscope views as a shared anchor to bridge surgical narrations and intraoperative OCT (iOCT) without requiring a fully synchronized tri-modal dataset, leveraging real microscope-narration videos and a synthetic dataset of synchronized microscope video and tool-aligned iOCT pairs. Contrastive alignment transfers structural priors from the synthetic domain to real videos lacking iOCT, and a dual-head MS-TCN++ integrates the resulting embeddings for joint macro- and micro-phase prediction. Evaluated on real vitreoretinal surgeries, our framework improves macro-phase recognition over a zero-shot baseline (mean F1 0.38 to 0.53) and provides an exploratory route to estimating fine-grained instrument-tissue measurements that are not directly observable in real microscope video alone; these micro-phase estimates are validated quantitatively on synthetic data and shown only qualitatively on real surgery. To our knowledge, this is the first work to unify microscope view, iOCT B-scans, and surgical narrations in a shared latent space for surgical phase recognition.
- Abstract(参考訳): 外科的位相認識は、硝子体手術における文脈認識型コンピュータ支援フィードバックの鍵となるが、同期型マルチモーダル術中データ(特に顕微鏡画像と術中CT)の不足は、多モーダル統合外科医を自然に再現することを目的としたアプローチを制限している。
対照的に、外科的ナレーションはオンラインで豊富に利用可能であり、豊富な意味的監督を提供する。
それまでの研究は主に対面でのコントラスト学習(例えば、手術中のOCT顕微鏡、顕微鏡ナレーション)を探求し、これら3つのモードの合同モデリングは、ほとんど探索されていない。
手術用ナレーションと手術用OCT(iOCT)を橋渡しするための共有アンカーとして顕微鏡ビューを用いたフレームワークを,完全な同期型トリモーダルデータセットを必要とせず,実際の顕微鏡ナレーションビデオと,同期型顕微鏡ビデオとツール対応のiOCTペアの合成データセットを活用する。
対照的に、コントラストアライメントは、合成ドメインからiOCTを欠いた実ビデオへ構造的先行を転送し、デュアルヘッドMS-TCN++は、結果として生じるマクロおよびマイクロフェーズ予測の埋め込みを統合する。
実際の硝子体手術では,ゼロショットベースライン(F1 0.38~0.53)のマクロ位相認識が向上し,顕微鏡ビデオのみで直接観察できない微細な計測機器を推定するための探索ルートが提供される。
我々の知る限り、これは顕微鏡視, iOCT B-scans, and surgery narrationsを外科的位相認識のための共有潜在空間に統一する最初の試みである。
関連論文リスト
- Monocular Marker-free Patient-to-Image Intraoperative Registration for Cochlear Implant Surgery [4.250558597144547]
本フレームワークは単分子型外科顕微鏡とシームレスに統合され,追加のハードウェア依存や要件を伴わずに臨床応用に極めて有用である。
以上より, 術中CTスキャンを2次元の手術シーンに登録し, 角誤差が10度以内の症例では, 6次元カメラの撮影位置の予測に臨床的に関連性があることが示唆された。
論文 参考訳(メタデータ) (2025-05-23T21:15:00Z) - Post-mastoidectomy Surface Multi-View Synthesis from a Single Microscopy Image [4.777201894011511]
単一CI顕微鏡画像から合成多視点映像を生成することができる新しいパイプラインを提案する。
本研究は, 術前CT検査を用いて, 乳頭切除後の表面を予測し, 本目的のために設計した方法である。
論文 参考訳(メタデータ) (2024-08-31T16:45:24Z) - CathFlow: Self-Supervised Segmentation of Catheters in Interventional Ultrasound Using Optical Flow and Transformers [66.15847237150909]
縦型超音波画像におけるカテーテルのセグメンテーションのための自己教師型ディープラーニングアーキテクチャを提案する。
ネットワークアーキテクチャは、Attention in Attentionメカニズムで構築されたセグメンテーショントランスフォーマであるAiAReSeg上に構築されている。
我々は,シリコンオルタファントムから収集した合成データと画像からなる実験データセット上で,我々のモデルを検証した。
論文 参考訳(メタデータ) (2024-03-21T15:13:36Z) - Monocular Microscope to CT Registration using Pose Estimation of the
Incus for Augmented Reality Cochlear Implant Surgery [3.8909273404657556]
本研究では, 外部追跡装置を必要とせず, 2次元から3次元の観察顕微鏡映像を直接CTスキャンに登録する手法を開発した。
その結果, x, y, z軸の平均回転誤差は25度未満, 翻訳誤差は2mm, 3mm, 0.55%であった。
論文 参考訳(メタデータ) (2024-03-12T00:26:08Z) - Style transfer between Microscopy and Magnetic Resonance Imaging via Generative Adversarial Network in small sample size settings [45.62331048595689]
磁気共鳴イメージング(MRI)のクロスモーダル増強と、同じ組織サンプルに基づく顕微鏡イメージングが期待できる。
コンディショナル・ジェネレーティブ・ディベサール・ネットワーク(cGAN)アーキテクチャを用いて,ヒト・コーパス・カロサムのMRI画像から顕微鏡組織像を生成する方法を検討した。
論文 参考訳(メタデータ) (2023-10-16T13:58:53Z) - Phase-Specific Augmented Reality Guidance for Microscopic Cataract
Surgery Using Long-Short Spatiotemporal Aggregation Transformer [14.568834378003707]
乳化白内障手術(英: Phaemulsification cataract surgery, PCS)は、外科顕微鏡を用いた外科手術である。
PCS誘導システムは、手術用顕微鏡映像から貴重な情報を抽出し、熟練度を高める。
既存のPCSガイダンスシステムでは、位相特異なガイダンスに悩まされ、冗長な視覚情報に繋がる。
本稿では,認識された手術段階に対応するAR情報を提供する,新しい位相特異的拡張現実(AR)誘導システムを提案する。
論文 参考訳(メタデータ) (2023-09-11T02:56:56Z) - CholecTriplet2021: A benchmark challenge for surgical action triplet
recognition [66.51610049869393]
腹腔鏡下手術における三肢の認識のためにMICCAI 2021で実施した内視鏡的視力障害であるColecTriplet 2021を提案する。
課題の参加者が提案する最先端の深層学習手法の課題設定と評価について述べる。
4つのベースライン法と19の新しいディープラーニングアルゴリズムが提示され、手術ビデオから直接手術行動三重項を認識し、平均平均精度(mAP)は4.2%から38.1%である。
論文 参考訳(メタデータ) (2022-04-10T18:51:55Z) - Multimodal Semantic Scene Graphs for Holistic Modeling of Surgical
Procedures [70.69948035469467]
カメラビューから3Dグラフを生成するための最新のコンピュータビジョン手法を利用する。
次に,手術手順の象徴的,意味的表現を統一することを目的としたマルチモーダルセマンティックグラフシーン(MSSG)を紹介する。
論文 参考訳(メタデータ) (2021-06-09T14:35:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。