論文の概要: HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
- arxiv url: http://arxiv.org/abs/2604.05961v1
- Date: Tue, 07 Apr 2026 14:55:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-08 17:42:09.895113
- Title: HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
- Title(参考訳): HumANDiff:モーション・コンセント・ヒューマン・ビデオ・ジェネレーションのための人工ノイズ拡散
- Authors: Tao Hu, Varun Jampani,
- Abstract要約: HumANDiffは、音声ノイズサンプリングによるビデオ拡散モデルを微調整することで、人間のビデオ生成を可能にする。
動作に一貫性があり、多様な服装スタイルを持つ高忠実な人間を実現している。
- 参考スコア(独自算出の注目度): 47.59225421284659
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Despite tremendous recent progress in human video generation, generative video diffusion models still struggle to capture the dynamics and physics of human motions faithfully. In this paper, we propose a new framework for human video generation, HumANDiff, which enhances the human motion control with three key designs: 1) Articulated motion-consistent noise sampling that correlates the spatiotemporal distribution of latent noise and replaces the unstructured random Gaussian noise with 3D articulated noise sampled on the dense surface manifold of a statistical human body template. It inherits body topology priors for spatially and temporally consistent noise sampling. 2) Joint appearance-motion learning that enhances the standard training objective of video diffusion models by jointly predicting pixel appearances and corresponding physical motions from the articulated noises. It enables high-fidelity human video synthesis, e.g., capturing motion-dependent clothing wrinkles. 3) Geometric motion consistency learning that enforces physical motion consistency across frames via a novel geometric motion consistency loss defined in the articulated noise space. HumANDiff enables scalable controllable human video generation by fine-tuning video diffusion models with articulated noise sampling. Consequently, our method is agnostic to diffusion model design, and requires no modifications to the model architecture. During inference, HumANDiff enables image-to-video generation within a single framework, achieving intrinsic motion control without requiring additional motion modules. Extensive experiments demonstrate that our method achieves state-of-the-art performance in rendering motion-consistent, high-fidelity humans with diverse clothing styles. Project page: https://taohuumd.github.io/projects/HumANDiff/
- Abstract(参考訳): 人間の動画生成の進歩にもかかわらず、生成的ビデオ拡散モデルは、人間の動きの力学と物理を忠実に捉えるのに苦戦している。
本稿では,人間の動画生成のための新しいフレームワークであるHumANDiffを提案する。
1) 統計的人体テンプレートの高密度表面多様体上にサンプリングされた非構造ランダムガウス雑音と非構造ランダムガウス雑音の時空間分布とを相関させた有声運動一貫性雑音サンプリングを行った。
空間的および時間的に一貫したノイズサンプリングのために、身体トポロジーを継承する。
2) 映像拡散モデルの標準訓練目標を高める共同運動学習は, 調音音から画素の出現と対応する身体運動を共同予測することで行う。
これは、例えば、動きに依存した衣服のしわをキャプチャする、高忠実な人間のビデオ合成を可能にする。
3) 音場に定義された新しい幾何運動整合性損失を用いて, フレーム間の物理運動整合性を強制する幾何学的運動整合性学習を行う。
HumANDiffは、音声によるノイズサンプリングによるビデオ拡散モデルを微調整することで、スケーラブルに制御可能な人間のビデオ生成を可能にする。
したがって,本手法は拡散モデル設計に非依存であり,モデルアーキテクチャの変更は不要である。
推論中、HumANDiffは単一のフレームワーク内で画像とビデオの生成を可能にし、追加のモーションモジュールを必要とせずに本質的なモーション制御を実現する。
広汎な実験により,動作に一貫性のある,多彩な衣服スタイルの高忠実な人体をレンダリングすることで,最先端の性能を実現することができた。
プロジェクトページ: https://taohuumd.github.io/projects/HumANDiff/
関連論文リスト
- MOSPA: Human Motion Generation Driven by Spatial Audio [83.31594478750682]
本稿では,多種多様で高品質な空間音声・動きデータを含む,空間音声駆動型人体運動データセットについて紹介する。
本研究では,身体運動と空間音声の関係を忠実に把握する,MOSPAと呼ばれるスパティアルオーディオによって駆動される人間の運動生成のためのフレームワークを開発する。
本手法は,本課題における最先端性能を実現する。
論文 参考訳(メタデータ) (2025-07-16T06:33:11Z) - M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation [65.48046909056468]
我々は,音声音声生成をビデオ前処理,モーション表現,レンダリング再構成を含む統一的なフレームワークに再構成する。
M2DAO-Talkerは2.43dBのPSNRの改善とユーザ評価ビデオの画質0.64アップで最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2025-07-11T04:48:12Z) - Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency [15.841490425454344]
本稿では,Loopy という,エンドツーエンドの音声のみの条件付きビデオ拡散モデルを提案する。
具体的には,ループ内時間モジュールとオーディオ・トゥ・ラテントモジュールを設計し,長期動作情報を活用する。
論文 参考訳(メタデータ) (2024-09-04T11:55:14Z) - Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation [5.78796187123888]
本研究では,1つの画像,ビデオ,ランダムノイズから時間的コヒーレントな人間のアニメーションを生成する手法を提案する。
両方向の時間的モデリングは、人間の外見の運動あいまいさを大幅に抑制することにより、生成ネットワーク上の時間的コヒーレンスを強制すると主張している。
論文 参考訳(メタデータ) (2023-07-02T13:57:45Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。