Fugu-MT 論文翻訳(概要): Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization

論文の概要: Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization

arxiv url: http://arxiv.org/abs/2606.10743v1
Date: Tue, 09 Jun 2026 11:53:29 GMT
ステータス: 翻訳完了
システム内更新日: 2026-06-11 16:42:38.029428
Title: Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization
Title（参考訳）: オープンワールド・コンタクト・ローカライゼーションによるビデオデモから手中心型人間-ロボット軌道移動
Authors: Yitian Shi, Di Wen, Zhengqi Han, Zicheng Guo, Yu Hu, Edgar Welte, Kunyu Peng, Rainer Stiefelhagen, Rania Rayyes,
Abstract要約: EmphHOWTransferは、人間のデモを接触認識、分類情報、多様なロボット軌道に蒸留する手中心のフレームワークである。 emphHOWTransferは、時間的に一貫した3次元手の動きを回復し、観察された手と物体の相互作用の手がかりを解析することで、時間的接触間隔を局所化する。実験によると、emphHOWTransferは86%の精度で正確な接触位置決めと高品質なロボットの動きを可能にする。
参考スコア（独自算出の注目度）: 28.23926683554352
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Learning from human video demonstrations remains challenging due to noisy hand-object interactions, unseen objects with partial observation, and cross-embodiment discrepancy. To address these challenges, we present \textit{HOWTransfer} (\emph{H}and-\emph{O}bject \emph{O}pen-\emph{W}orld Transfer), a hand-centric framework that distills human demonstrations into contact-aware, taxonomy-informed, and diverse robotic trajectories. Instead of relying on object-specific descriptions, vision-language queries, or explicit object-state tracking, \emph{HOWTransfer} recovers temporally consistent 3D hand motion and localizes temporal contact intervals by reasoning over observed hand-object interaction cues. The localized contact onsets are then used to retarget human grasp intent into multi-modal parallel-jaw grasp hypotheses, which are propagated along the recovered wrist trajectory to generate robot-executable motions. Finally, a trajectory editing stage refines contact alignment and produces diverse executable variants from a single demonstration. Experiments across diverse manipulation tasks show that \emph{HOWTransfer} enables accurate contact localization and high-quality robot motion retargeting with $86\%$ success, which is preferred over teleoperated trajectories in a blinded preference study.
Abstract（参考訳）: 人間のビデオのデモから学ぶことは、ノイズの多い手-物体の相互作用、部分的な観察を伴う見えない物体、異体間不一致など、依然として困難である。これらの課題に対処するために,人間の実演を接触認識,分類情報,多種多様なロボット軌道に蒸留する手中心のフレームワークであるtextit{HOWTransfer} (\emph{H}and-\emph{O}bject \emph{O}pen-\emph{W}orld Transfer)を提案する。オブジェクト固有の記述、視覚言語クエリ、明示的なオブジェクト状態追跡に頼る代わりに、 \emph{HOWTransfer} は時間的に一貫した3次元手の動きを回復し、観察された手と物体の相互作用の手がかりを引き合いに出して時間的接触間隔を局所化する。局所化されたコンタクトオンセットは、人間のつかむ意図をマルチモーダルなパラレルジャウグリップ仮説に再ターゲティングするために使用され、この仮説は、回復した手首軌道に沿って伝播して、ロボットが実行可能な動作を生成する。最後に、軌道編集段階は、接触アライメントを洗練させ、単一のデモンストレーションから様々な実行可能な変種を生成する。多様な操作タスクを対象とした実験では, 視覚障害者を対象にした遠隔操作よりも, 正確な接触位置決めと高品質なロボット動作のリターゲティングを実現している。

論文の概要: Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization

関連論文リスト