論文の概要: Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
- arxiv url: http://arxiv.org/abs/2607.02798v1
- Date: Thu, 02 Jul 2026 22:17:20 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.422586
- Title: Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
- Title(参考訳): 3D-Grounded Motion-Consistent Noise for Controllable Video Generation (特集:3D-Grounded Motion-Consistent Noise)
- Authors: Long Vu, Tan Ngo, Animesh Karnewar, Amir Habibian, Binh-Son Hua, Hung Bui, Minh Hoai Nguyen, Phong Nguyen-Ha,
- Abstract要約: オブジェクト軌跡とカメラ視点の同時制御を可能にする統合フレームワークUniCaMoを提案する。
UniCaMoは、潜伏するビデオフレームにまたがる3Dグラウンドのモーションコンシステントノイズスペースを共有できる。
標準的な制御可能なビデオ生成ベンチマークにおいて、映像品質とモーション制御性の両方を最新技術で実現している。
- 参考スコア(独自算出の注目度): 14.852131707037655
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo builds a shared 3D-grounded motion-consistent noise space across latent video frames. Sparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global sphere-based noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.
- Abstract(参考訳): 現代の画像とテキストとビデオの拡散モデルは、参照画像とテキスト入力に条件付けられた初期ガウスノイズテンソルを反復的にデノベートすることで、非常にリアルなビデオを合成することができる。
しかし、既存のアプローチは、単一生成プロセスにおいて、オブジェクトの動きとカメラの動きの両方に対して、正確かつ統一的な制御性を欠いている。
拡散モデルの入力ノイズを直接構成することにより、オブジェクト軌跡とカメラ視点の同時制御を可能にする統合フレームワークUniCaMoを提案する。
具体的には、UniCaMoは3Dグラウンドのモーションコンセントのノイズ空間を、潜伏するビデオフレーム間で共有する。
スパース3Dポイントトラックは、所望の物体軌跡に沿って基準フレームのガウスノイズをワープするために使用され、一方、仮想球面ノイズ表現は、カメラモーションの下で新たに明らかになったシーン領域に対して、一様に一貫したノイズ値を提供する。
局所トラック誘導ノイズワープとグローバルスフィアベースノイズサンプリングを組み合わせることで、UniCaMoは物体の動きと視点の変化の両方の下で幾何的および時間的一貫性を維持する。
UniCaMoは入力ノイズのみを変更するため、補助アダプタ、制御ブランチ、基礎となるビデオ拡散モデルへのアーキテクチャ変更は不要である。
Wan 2.1 (14B)を含む大規模ビデオ拡散モデルの軽量なLoRA微調整により、UniCaMoは、標準的な制御可能なビデオ生成ベンチマークにおいて、ビデオ品質とモーション制御性の両方において最先端の結果を達成する。
関連論文リスト
- CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping [81.58057192696843]
数値カメラパラメータを直接拡散バックボーンに注入する既存の方法は、抽象座標と視覚的内容の間のギャップを埋めることがしばしば失敗する。
本稿では,カメラの動きを時間的コヒーレントな表現に符号化するフロー・ツー・ノイズ・ワープ手法であるCameraNoiseを提案する。
CameraNoiseを拡散プロセスに統合することで、我々のフレームワークは安定した高忠実度ビデオを提供する。
論文 参考訳(メタデータ) (2026-05-29T03:02:50Z) - HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation [47.59225421284659]
HumANDiffは、音声ノイズサンプリングによるビデオ拡散モデルを微調整することで、人間のビデオ生成を可能にする。
動作に一貫性があり、多様な服装スタイルを持つ高忠実な人間を実現している。
論文 参考訳(メタデータ) (2026-04-07T14:55:10Z) - M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation [65.48046909056468]
我々は,音声音声生成をビデオ前処理,モーション表現,レンダリング再構成を含む統一的なフレームワークに再構成する。
M2DAO-Talkerは2.43dBのPSNRの改善とユーザ評価ビデオの画質0.64アップで最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2025-07-11T04:48:12Z) - Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise [19.422355461775343]
我々は、構造化潜在雑音サンプリングによる動き制御を可能とし、映像拡散モデルを強化した。
本稿では,ランダムな時空間のガウス性と相関した雑音を置き換え,リアルタイムに動作可能な新しいノイズワープアルゴリズムを提案する。
提案アルゴリズムの効率性により,ワープノイズを最小限のオーバーヘッドで使用することで,最新の映像拡散ベースモデルを微調整することができる。
論文 参考訳(メタデータ) (2025-01-14T18:59:10Z) - Learning Task-Oriented Flows to Mutually Guide Feature Alignment in
Synthesized and Real Video Denoising [137.5080784570804]
Video Denoisingは、クリーンなノイズを回復するためにビデオからノイズを取り除くことを目的としている。
既存の研究によっては、近辺のフレームから追加の空間的時間的手がかりを利用することで、光学的流れがノイズ発生の助けとなることが示されている。
本稿では,様々なノイズレベルに対してより堅牢なマルチスケール光フロー誘導型ビデオデノイング法を提案する。
論文 参考訳(メタデータ) (2022-08-25T00:09:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。