Fugu-MT 論文翻訳(概要): Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

論文の概要: Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

arxiv url: http://arxiv.org/abs/2605.22272v2
Date: Fri, 22 May 2026 04:19:09 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-25 14:44:53.777109
Title: Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
Title（参考訳）: imagine2Real: ビデオ生成プリミティブによるゼロショットヒューマノイドオブジェクトインタラクションを目指して
Authors: Jiahe Chen, ZiRui Wang, Feiyu Jia, Xiao Chen, Xiaojie Niu, Weishuai Zeng, Tianfan Xue, Xiaowei Zhou, Jiangmiao Pang, Jingbo Wang,
Abstract要約: 高忠実度3Dデータの不足により,全体Humanoid-Object Interaction (HOI) がボトルネックとなる。本研究では,ゼロショットHOIフレームワークであるImagine2Realを提案する。
参考スコア（独自算出の注目度）: 51.096845970243855
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to their reliance on geometric priors (e.g., explicit CAD models), and \textit{Retargeting Complexity} arising from intensive morphing and morphological mismatch. We propose Imagine2Real, a zero-shot HOI framework for flexible, geometry-free interaction. To resolve misalignment, we formulate robot and object motions as unified 4D point trajectories. To overcome retargeting complexity, our Keypoints Tracker tracks only sparse critical points (base, hands, and object), entirely bypassing the error-amplifying retargeting process. To maintain natural gaits despite these sparse signals, we utilize the latent space of a Behavior Foundation Model (BFM) as the tracker's search domain. Using a progressive training strategy, Imagine2Real learns robust behaviors with simple tracking rewards, enabling zero-shot physical deployment within a motion capture(mocap) system.
Abstract（参考訳）: 高忠実度3Dデータの不足により,全体Humanoid-Object Interaction (HOI) がボトルネックとなる。ビデオ生成の先行は有望な代替手段を提供するが、既存の手法は、幾何学的先行(例えば、明示的なCADモデル)と、集中的なモルヒネや形態的ミスマッチから生じる‘textit{Retargeting Complexity’に依存するため、 'textit{Representation Misalignment' に苦しむ。本研究では,ゼロショットHOIフレームワークであるImagine2Realを提案する。誤認識を解決するため,ロボットと物体の動きを統合された4次元点軌道として定式化する。再ターゲティングの複雑さを克服するために、Keypoints Trackerは、エラーを増幅する再ターゲティングプロセスを完全にバイパスする、わずかなクリティカルポイント(ベース、ハンド、オブジェクト)のみをトラックします。これらの疎い信号にもかかわらず、自然視線を維持するために、トラッカーの探索領域として振舞い基礎モデル(BFM)の潜在空間を利用する。プログレッシブトレーニング戦略を使用して、Imagine2Realは単純なトラッキング報酬で堅牢な動作を学び、モーションキャプチャ(mocap)システム内でゼロショットの物理的なデプロイメントを可能にする。

論文の概要: Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

関連論文リスト