論文の概要: WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
- arxiv url: http://arxiv.org/abs/2608.15659v2
- Date: Wed, 19 Aug 2026 05:36:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-20 13:35:39.174911
- Title: WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
- Title(参考訳): WorldRover: リッチアノテーションを備えた世界探索のためのスケーラブルな合成ビデオデータエンジン
- Abstract要約: 我々は、アーティストが構築した環境のリッチな注釈付き長距離探索を生成するデータエンジンであるWorldRoverを紹介した。
WorldRover-EngineはUnreal Engineパイプラインであり、ミニスケールのルートを実行およびオフラインで実行する。
さらに3人目のサブセットは、高密度な光学フロー、可視性を持つ長距離2D/3Dポイントトラック、およびカメラの軌跡とは異なるキャラクタ軌道を提供する。
- 参考スコア(独自算出の注目度): 29.798429925680782
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
- Abstract(参考訳): 探索可能な世界を生成または再構成するためには、カメラモーション、シーン幾何学、時間対応、インタラクティブなモデル、制御信号など、RGB以上のビデオを組み合わせる必要がある。
リアルキャプチャーはこれらの信号のいくつかを提供することができるが、密度の高い幾何学と長距離対応は通常、推定や特殊化された計器に依存している。
レンダリングはこれらの量を直接提供しますが、既存の合成資源が同じフレーム上に組み合わされることは滅多にありません。
我々は、アーティストが構築した環境のリッチな注釈付き長距離探索を生成するデータエンジンであるWorldRoverを紹介した。
中心となるWorldRover-EngineはUnreal Engineパイプラインであり、完全な軌跡とシーン幾何学を保ちながら、ミニスケールのルートを実行およびオフラインで実行する。
同じ探査は、環境条件の異なる一対一、三対三、360パノラマカメラから行うことができる。
我々はWorldRover-Engineを用いて,RGBと距離深度,カメラ軌跡,軌跡由来の動作信号とをそれぞれ組み合わせたWorldRover-10Mを構築した。
さらに3人目のサブセットは、高密度な光学フロー、可視性を持つ長距離2D/3Dポイントトラック、およびカメラの軌跡とは異なるキャラクタ軌道を提供する。
エンジンは、ルートとシーンの幾何学を保ちながら、異なる環境状態または中立な白色物質の下で、一人称、三人称、360パノラマ的な視点からトラバーサルを描画することができる。
したがって、WorldRoverは、長期的世界探査をスケーラブルなデータ生成問題に転換し、探索可能な世界の一貫性のある表現を構築し、維持し、再考しなければならないモデルの監督を提供する。
関連論文リスト
- RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction [38.41034045192577]
フィードフォワード3次元再構成の最近の進歩は、画像ストリームのみから高密度なシーン表現とカメラモーションを復元できるモデルを生み出した。
重力整列全方位ビデオのための大規模再構成パイプラインRIGORを提案する。
提案した整合性機構は, 建設現場の挑戦的配列に基づいて, フィードフォワードベースライン上での軌道精度と再構成幾何の両方を改善することを実証する。
論文 参考訳(メタデータ) (2026-09-11T20:11:43Z) - ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space [49.372892542771886]
ABot-3DWorld 0は、テキスト、画像、ビデオの入力を高忠実で探索可能な3D世界に変換する、普遍的なマルチモーダルな3Dワールドモデルである。
論文 参考訳(メタデータ) (2026-07-13T15:14:48Z) - Lyra 2.0: Explorable Generative 3D Worlds [77.45279013687427]
Lyra 2.0は、永続的で探索可能な3D世界を大規模に生成するためのフレームワークです。
空間的忘れに対処するため、フレームごとの3D形状を維持し、情報ルーティングのみに使用します。
自己拡張された履歴をトレーニングして、モデルを自身の劣化した出力に公開し、それを伝播するのではなく、ドリフトを正すように教えます。
論文 参考訳(メタデータ) (2026-04-14T17:59:44Z) - SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations [88.8747004592363]
SceneScribe-1Mは新しい大規模多時間ビデオデータセットである。
そこには100万本のビデオが含まれており、それぞれに詳細なテキスト記述、正確なパラメータ、深度マップ、一貫性のある3Dポイントトラックなどが含まれている。
SceneScribe-1Mの汎用性と価値は、単眼深度推定、シーン再構成動的点追跡、テキスト・ビデオ合成などの生成タスク、カメラ制御の有無にかかわらず、幅広い下流タスクのベンチマークを確立することで示される。
論文 参考訳(メタデータ) (2026-04-09T08:59:33Z) - WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling [63.37476802589492]
We present WorldReel, a 4D video that are native-temporally consistent。
WorldReelは、ポイントマップ、カメラ軌道、高密度フローを含む4Dシーン表現と共にフレームを生成する。
We believe that WorldReel bring video generation to 4D-consistent world modeling。
論文 参考訳(メタデータ) (2025-12-08T18:54:12Z) - Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation [66.95956271144982]
本稿では,単一画像から一貫した3Dポイントクラウドシーケンスを生成する新しいビデオ拡散フレームワークであるVoyagerを紹介する。
既存のアプローチとは異なり、Voyagerはフレーム間で固有の一貫性を持って、エンドツーエンドのシーン生成と再構築を実現している。
論文 参考訳(メタデータ) (2025-06-04T17:59:04Z) - WorldExplorer: Towards Generating Fully Navigable 3D Scenes [48.16064304951891]
WorldExplorerは、幅広い視点で一貫した視覚的品質で、完全にナビゲート可能な3Dシーンを構築する。
私たちは、シーンを深く探求する、短く定義された軌道に沿って、複数のビデオを生成します。
我々の新しいシーン記憶は、各ビデオが最も関連性の高い先行ビューで条件付けされている一方、衝突検出機構は劣化を防止している。
論文 参考訳(メタデータ) (2025-06-02T15:41:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。