Fugu-MT 論文翻訳(概要): SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning

論文の概要: SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning

arxiv url: http://arxiv.org/abs/2512.18068v1
Date: Fri, 19 Dec 2025 21:15:26 GMT
ステータス: 翻訳完了
システム内更新日: 2026-03-23 08:17:40.439812
Title: SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
Title（参考訳）: SurgiPose:手術用ロボット学習のための単眼ビデオから手術用ツールキネマティクスを推定する
Authors: Juo-Tung Chen, XinHao Chen, Ji Woong Kim, Paul Maria Scheikl, Richard Jaepyeong Cha, Axel Krieger,
Abstract要約: 大規模な外科的デモンストレーションの有望な情報源は、モノラルな外科的ビデオがオンラインで公開されていることである。本稿では,単眼手術映像から運動情報を推定するための,微分レンダリングに基づくアプローチであるSurgiPoseを提案する。本研究の結果から, 推定キネマティクスに基づいてトレーニングした政策は, 真理データに基づいてトレーニングした政策に匹敵する成功率を示した。
参考スコア（独自算出の注目度）: 5.304104136888954
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Imitation learning (IL) has shown immense promise in enabling autonomous dexterous manipulation, including learning surgical tasks. To fully unlock the potential of IL for surgery, access to clinical datasets is needed, which unfortunately lack the kinematic data required for current IL approaches. A promising source of large-scale surgical demonstrations is monocular surgical videos available online, making monocular pose estimation a crucial step toward enabling large-scale robot learning. Toward this end, we propose SurgiPose, a differentiable rendering based approach to estimate kinematic information from monocular surgical videos, eliminating the need for direct access to ground truth kinematics. Our method infers tool trajectories and joint angles by optimizing tool pose parameters to minimize the discrepancy between rendered and real images. To evaluate the effectiveness of our approach, we conduct experiments on two robotic surgical tasks: tissue lifting and needle pickup, using the da Vinci Research Kit Si (dVRK Si). We train imitation learning policies with both ground truth measured kinematics and estimated kinematics from video and compare their performance. Our results show that policies trained on estimated kinematics achieve comparable success rates to those trained on ground truth data, demonstrating the feasibility of using monocular video based kinematic estimation for surgical robot learning. By enabling kinematic estimation from monocular surgical videos, our work lays the foundation for large scale learning of autonomous surgical policies from online surgical data.
Abstract（参考訳）: イミテーション・ラーニング(IL)は、外科的タスクの学習を含む、自律的な器用な操作を可能にするという大きな可能性を示している。手術におけるILの可能性を完全に解放するには、臨床データセットへのアクセスが必要であるが、残念ながら現在のILアプローチに必要なキネマティックなデータは欠如している。大規模な外科的デモンストレーションの有望な情報源は、モノラルな手術ビデオがオンラインで公開されているため、モノラルなポーズ推定が大規模なロボット学習を実現するための重要なステップとなる。この目的に向けて,単眼手術ビデオからキネマティック情報を推定する,微分レンダリングに基づくアプローチであるSurgiPoseを提案し,真理キネマティックスへの直接アクセスの必要性を排除した。本手法は,レンダリング画像と実画像との差を最小限に抑えるために,ツールポーズパラメータを最適化することにより,ツール軌跡と関節角度を推定する。提案手法の有効性を評価するため,da Vinci Research Kit Si (dVRK Si) を用いて, 組織リフトとニードルピックアップの2つのロボット外科的課題について実験を行った。本研究では,実測値と実測値の両方をビデオから学習し,実測値と実測値とを比較した。以上の結果から, 単眼映像による体力評価を手術ロボット学習に応用できる可能性が示唆された。単眼手術ビデオからキネマティックな評価を可能にすることで,オンライン手術データから自律的な手術方針を大規模に学習する基盤を構築した。

論文の概要: SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning

関連論文リスト