論文の概要: World2Motion: Turning Video World Models into 3D Human Motion Generators
- arxiv url: http://arxiv.org/abs/2609.37004v2
- Date: Wed, 30 Sep 2026 12:19:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:26.210215
- Title: World2Motion: Turning Video World Models into 3D Human Motion Generators
- Title(参考訳): World2Motion:ビデオの世界モデルを3Dのモーションジェネレータに変える
- Abstract要約: We present World2Motion, a framework that a scene-aware human motion and corresponding video from a single image and a text prompt。
この適応には、ペアビデオモーションデータの不足と、生成された動きにおける時間的不安定という2つの課題がある。
マルチソースインタラクションベンチマークの実験では、World2Motionは評価された3Dモーションジェネレータと比較して、テキストアライメントとシーンインタラクションが優れていることが示されている。
- 参考スコア(独自算出の注目度): 33.66283404630084
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference. Our project page is available at https://fyantu.github.io/World2Motion/.
- Abstract(参考訳): We present World2Motion, a framework that a scene-aware human motion and corresponding video from a single image and a text prompt。
既存の3Dモーションジェネレータはモーションデータセットから学習するが、その一般化は環境の限られた範囲によって制限される。
対照的に、コスモス3のようなビデオワールドモデルはより広い環境条件を提供するが、フルボディのモーション生成には設計されていない。
これらに対処するため、コスモス3を1段の3Dモーションジェネレータに変える。
この適応には2つの課題がある: ペアビデオモーションデータの不足と生成された動きの時間的不安定性。まず、合成ビデオモーションペアと推定3Dモーションのペアを組み合わせた実ビデオを組み合わせたトレーニングデータセットを構築する。
第2に,異なるノイズレベルをビデオとモーションに割り当てるシフト分離型ノイズスケジュールを提案する。
この設計は2つのモードの異なる復調要求に対応し、モーションジッタを低減させる。
マルチソースインタラクションベンチマークの実験では、World2Motionは評価された3Dモーションジェネレータと比較して、テキストアライメントとシーンインタラクションが優れていることが示されている。
また、2段階のベースラインの相互作用の成功率と一致し、およそ3.3$\times$高速推論を達成した。
私たちのプロジェクトページはhttps://fyantu.github.io/World2Motion/.comで公開されている。
関連論文リスト
- DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data [51.316274891736164]
DynaVidは、トレーニングで合成モーションデータを活用するビデオ合成フレームワークである。
ダイナミックモーション生成とカメラモーション制御において,DynaVidはリアリズムと制御性を向上することを示す。
論文 参考訳(メタデータ) (2026-04-02T06:12:38Z) - VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models [110.32291962407078]
VimoRAG(ヴィモラグ)は、動画に基づく大規模言語モデルのためのモーション生成フレームワークである。
動作中心の効果的なビデオ検索モデルを開発し、最適下検索結果による誤り伝播の問題を緩和する。
実験結果から,VimoRAGはテキストのみの入力に制約された動きLLMの性能を大幅に向上させることがわかった。
論文 参考訳(メタデータ) (2025-08-16T15:31:14Z) - M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation [65.48046909056468]
我々は,音声音声生成をビデオ前処理,モーション表現,レンダリング再構成を含む統一的なフレームワークに再構成する。
M2DAO-Talkerは2.43dBのPSNRの改善とユーザ評価ビデオの画質0.64アップで最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2025-07-11T04:48:12Z) - Motion-2-to-3: Leveraging 2D Motion Data to Boost 3D Motion Generation [43.915871360698546]
人間の2Dビデオは、幅広いスタイルやアクティビティをカバーし、広範にアクセス可能なモーションデータのソースを提供する。
本研究では,局所的な関節運動をグローバルな動きから切り離し,局所的な動きを2次元データから効率的に学習する枠組みを提案する。
提案手法は,2次元データを効率的に利用し,リアルな3次元動作生成をサポートし,支援対象の動作範囲を拡大する。
論文 参考訳(メタデータ) (2024-12-17T17:34:52Z) - Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes [90.39860012099393]
Sitcom-Crafterは3D空間における人間のモーション生成システムである。
機能生成モジュールの中心は、我々の新しい3Dシーン対応ヒューマン・ヒューマン・インタラクションモジュールである。
拡張モジュールは、コマンド生成のためのプロット理解、異なるモーションタイプのシームレスな統合のためのモーション同期を含む。
論文 参考訳(メタデータ) (2024-10-14T17:56:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。