論文の概要: Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
- arxiv url: http://arxiv.org/abs/2608.28174v1
- Date: Fri, 28 Aug 2026 10:39:52 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-31 17:16:04.29121
- Title: Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
- Title(参考訳): Manifold4D: ビデオ再撮影のためのポイントクラウドレンダリングマニフォールド
- Authors: Yongqi Mao, Zijia Dai, Zhishuo Liu, Wei Xu, Kaiwei Wang, Guotao Meng,
- Abstract要約: ビデオの再撮影は、ユーザが指定したカメラ軌道に沿って、ダイナミックシーンのモノクロビデオを再レンダリングする。
支配的なレシピは、ターゲットの幾何学を明示的に提供します。
フローマッチングの初期ノイズに直接レンダリングを注入するMANIFOLD4Dを提案する。
- 参考スコア(独自算出の注目度): 16.163736801043942
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Video re-shooting re-renders a monocular video of a dynamic scene along a user-specified camera trajectory, and the dominant recipe supplies the target geometry explicitly: per-frame depth lifts the source video into a 4D point cloud, which is rasterized along the trajectory into a point cloud render. Because the render and the source video are both handed to the network as visual conditions, they compete at every denoising step, leaving the model with a trust dilemma --- how much of the render to believe --- which can degrade trajectory control or visual quality on data outside the training distribution. We argue that a render already pixel-aligned with the target view does not need to be supplied as an explicit conditioning stream at all. We propose MANIFOLD4D, which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a new noise manifold carrying geometric information, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our DAVIS-Traj benchmark and on the Vista4D evaluation set, MANIFOLD4D attains the best camera-control accuracy on every metric, lowering rotation error by 25% and 27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and leading on real-world novel-view photometric quality. In a user study, our method achieves clear advantages in trajectory following and dynamic consistency. The gap widens as the yaw amplitude grows past the training range, and the model still recovers correct dynamic motion from the source video when the render is deliberately corrupted, confirming that the geometric prior guides generation without overriding it.
- Abstract(参考訳): ビデオの再レンダリングは、ユーザが指定したカメラ軌道に沿って、ダイナミックシーンのモノラルなビデオを再レンダリングし、支配的なレシピは、ターゲットの形状を明示的に提供します。
レンダリングとソースビデオはどちらも視覚的な条件としてネットワークに渡されるため、トレーニングディストリビューション外のデータに対して、トラジェクティブコントロールや視覚的品質を低下させる可能性のある、信頼されたジレンマ -- をモデルに残して、すべての視覚的なステップで競合する。
我々は、既にターゲットビューに一致したレンダリングを明示的な条件付けストリームとして供給する必要はないと主張している。
フローマッチングの初期ノイズに直接レンダリングを注入するMANIFOLD4Dを提案する。これにより、生成は標準ガウスノイズから外れず、幾何学的情報を持つ新しいノイズ多様体から切り離され、ソースビデオが唯一の視覚条件となる。
したがって、レンダリングは正確に1回だけ使用され、ネットワークは読み方を学ぶように要求されることはない。
DAVIS-TrajベンチマークとVista4D評価セットでは、MANIFOLD4Dは、すべてのメートル法で最高のカメラ制御精度を達成し、ローテーションエラーを25%と27%、翻訳エラーを最大32%まで下げ、ビデオの忠実度でマッチングし、実世界のノベルビューの測光品質に導いた。
提案手法は, 軌道追従と動的整合性において明らかな優位性を実現する。
ヤウ振幅がトレーニング範囲を超えて大きくなるにつれてギャップは拡大し、レンダリングが意図的に破損したときにも、モデルが元の動画から正しいダイナミックな動きを回復し、幾何学的先行ガイドがオーバーライドすることなく生成することを確認した。
関連論文リスト
- Vista4D: Video Reshooting with 4D Point Clouds [60.22650781486302]
Vista4Dは、入力されたビデオとターゲットカメラを4Dポイントクラウドにグラウンドするビデオリシューティングフレームワークである。
現状のベースラインと比較して、4Dの一貫性、カメラ制御、視覚的品質が改善された。
本手法は,動的シーン展開や4次元シーン再構成といった実世界の応用に一般化する。
論文 参考訳(メタデータ) (2026-04-23T17:57:28Z) - Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting [3.1328424544428852]
インターネット規模のモノクロビデオを活用するためのフレームワークを構築した。
私たちのコアコントリビューションは、ソースビデオ、幾何アンカー、ターゲットビデオからなる擬似多視点トレーニング三脚の生成です。
提案する拡散変圧器は4Dポイントクラウド誘導アンカーを用いて,最先端の時間的整合性を実現する。
論文 参考訳(メタデータ) (2026-04-23T15:32:56Z) - LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models [52.656349227001925]
モノクロビデオが与えられた場合、ビデオの再レンダリングの目的は、新しいカメラの軌跡からシーンのビューを生成することである。
既存の方法は2つの異なる課題に直面している。
大規模な4次元再構成モデルの潜在空間に埋め込まれた暗黙的幾何学的知識を用いて,これらの課題に対処することを提案する。
論文 参考訳(メタデータ) (2026-01-21T05:46:03Z) - DriveExplorer: Images-Only Decoupled 4D Reconstruction with Progressive Restoration for Driving View Extrapolation [12.714087160353317]
本稿では,自律運転シナリオにおける視線外挿の有効解を提案する。
近年のアプローチは、拡散モデルを用いて、与えられた視点からシフトした新しいビュー画像を生成することに焦点を当てている。
本手法は,ベースラインと比較して,新規な外挿視点で高品質な画像を生成する。
論文 参考訳(メタデータ) (2025-12-30T04:41:56Z) - Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models [79.06910348413861]
Diff4Splatは、単一の画像から制御可能で明示的な4Dシーンを合成するフィードフォワード方式である。
単一の入力画像、カメラ軌跡、オプションのテキストプロンプトが与えられた場合、Diff4Splatは外見、幾何学、動きを符号化する変形可能な3Dガウス場を直接予測する。
論文 参考訳(メタデータ) (2025-11-01T11:16:25Z) - SEE4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting [83.5106058182799]
SEE4Dは, カジュアルビデオから4次元世界モデリングを行うための, ポーズのないトラジェクトリ・ツー・カメラ・フレームワークである。
モデル内のビュー条件ビデオは、現実的に合成された画像を認知する前に、ロバストな幾何学を学ぶために訓練される。
クロスビュービデオ生成とスパース再構成のベンチマークでSee4Dを検証した。
論文 参考訳(メタデータ) (2025-10-30T17:59:39Z) - Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video [56.781766315691854]
ビデオ条件付き4D再生のための幾何学保存パイプラインである textbfRestage4D を紹介する。
DAVIS と PointOdyssey 上のRestage4D の有効性を検証し,幾何整合性,運動品質,3次元追跡性能の向上を実証した。
論文 参考訳(メタデータ) (2025-08-08T21:31:51Z) - Deblur4DGS: 4D Gaussian Splatting from Blurry Monocular Video [55.704264233274294]
ぼやけたモノクロ映像から高品質な4Dモデルを再構成するためのDeblur4DGSを提案する。
我々は露光時間内の連続的動的表現を露光時間推定に変換する。
Deblur4DGSは、新規なビュー合成以外にも、複数の視点からぼやけたビデオを改善するために応用できる。
論文 参考訳(メタデータ) (2024-12-09T12:02:11Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。