論文の概要: Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
- arxiv url: http://arxiv.org/abs/2606.28828v1
- Date: Sat, 27 Jun 2026 09:27:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-30 18:07:15.700598
- Title: Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
- Title(参考訳): Ground4D:モノクロビデオからの一貫性を意識した4D再構成
- Authors: Qing Zhao, Weijian Deng, Pengxu Wei, Liang Lin,
- Abstract要約: 2つのステージ上に構築された幾何グラウンドのフレームワークであるGround4Dを提案する。
まず,モノクロ映像から複数視点の立体形状とカメラポーズを再構成する。
第2に,動的ガウススプラッティングによる幾何整合性を考慮した改良を行う。
- 参考スコア(独自算出の注目度): 69.94904173314505
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over time remains challenging. Dynamic Gaussian Splatting achieves strong rendering performance through photometric optimization, yet does not explicitly enforce multi-view geometric consistency. In contrast, 3D foundation models recover coherent scene geometry and camera motion, but their point-based outputs are not designed for photorealistic rendering. We propose Ground4D, a geometry-grounded framework built on two stages. First, we perform geometry initialization via 3D foundation models, leveraging VGGT in a training-free manner to reconstruct multi-view-consistent 3D geometry and camera poses from monocular video. The recovered geometry provides a structured and reliable initialization for dynamic Gaussian representations. Second, we conduct geometry-consistency-aware refinement via dynamic Gaussian Splatting, optimizing the representation through differentiable rendering while maintaining multi-view geometric consistency across both observed and synthesized viewpoints. Furthermore, Ground4D inherently models the continuous 4D dynamics of the scene, naturally supporting rendering at arbitrary timestamps. By integrating foundation-level geometric priors into dynamic Gaussian optimization, Ground4D achieves stronger reconstruction fidelity and rendering performance, underscoring the role of geometry-grounded constraints in robust 4D scene modeling.
- Abstract(参考訳): 動的ノベルビュー合成をサポートしながら、時間とともに忠実な幾何学を維持しながら、4Dシーンの表現を単一のモノクロビデオから学習することは、依然として困難である。
動的ガウススプラッティングは、光度最適化によって強いレンダリング性能を達成するが、多視点幾何一貫性を明示的に強制するわけではない。
対照的に、3Dファウンデーションモデルは、コヒーレントなシーン形状とカメラモーションを復元するが、ポイントベースの出力は、フォトリアリスティックなレンダリングのために設計されていない。
2つのステージ上に構築された幾何グラウンドのフレームワークであるGround4Dを提案する。
まず,VGGTをトレーニング不要な方法で活用し,モノクロ映像から多視点対応の3D幾何とカメラポーズを再構成する。
復元された幾何学は、動的ガウス表現に対して構造化され信頼性の高い初期化を与える。
第2に、動的ガウススプラッティングによる幾何整合性を考慮した改良を行い、観察された視点と合成された視点の多視点的整合性を維持しながら、微分可能なレンダリングによる表現を最適化する。
さらに、Ground4Dは本質的にシーンの連続した4Dダイナミクスをモデル化し、任意のタイムスタンプでのレンダリングを自然にサポートする。
基礎レベルの幾何学的前提を動的ガウスの最適化に組み込むことで、グラウンド4Dはより強力な再構成忠実度とレンダリング性能を実現し、ロバストな4次元シーンモデリングにおける幾何学的制約の役割を浮き彫りにする。
関連論文リスト
- IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation [76.36174247570716]
ポーズレス多視点画像から連続的かつ一貫性のある幾何を暗黙的にモデル化するインプリシトビジュアル幾何変換器IVGTを提案する。
IVGTは標準座標系で連続的なニューラルネットワークシーン表現を学習し、任意の3D位置での連続的な空間クエリをサポートする。
連続的かつコヒーレントな表面形状の直接抽出を可能にし、任意の視点からRGB画像、深度マップ、表面正規写像のレンダリングを可能にする。
論文 参考訳(メタデータ) (2026-05-15T17:59:57Z) - Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis [53.48281548500864]
Motion 3-to-4は、単一のモノクロビデオから高品質な4Dダイナミックオブジェクトを合成するためのフィードフォワードフレームワークである。
我々のモデルは、コンパクトな動き潜在表現を学習し、フレーム単位の軌道を予測して、時間的コヒーレントな幾何である完全なロバスト性を取り戻す。
論文 参考訳(メタデータ) (2026-01-20T18:59:48Z) - Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image [88.71287865590273]
そこでTrajScene-60Kについて紹介する。
拡散型4次元シーン軌道生成装置(4D-STraG)を提案する。
次に、4Dポイントトラック表現から任意のカメラトラジェクトリでビデオをレンダリングする4Dビュー合成モジュール(4D-Vi)を提案する。
論文 参考訳(メタデータ) (2025-12-04T17:59:10Z) - Can Video Diffusion Model Reconstruct 4D Geometry? [66.5454886982702]
Sora3Rは、カジュアルなビデオから4Dのポイントマップを推測するために、大きなダイナミックビデオ拡散モデルのリッチ・テンポラリなテンポラリなテンポラリな時間を利用する新しいフレームワークである。
実験により、Sora3Rはカメラのポーズと詳細なシーン形状の両方を確実に復元し、動的4D再構成のための最先端の手法と同等の性能を発揮することが示された。
論文 参考訳(メタデータ) (2025-03-27T01:44:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。