論文の概要: Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models
- arxiv url: http://arxiv.org/abs/2610.01614v1
- Date: Thu, 01 Oct 2026 12:52:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.136238
- Title: Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models
- Title(参考訳): Oneira: ビデオワールドモデルにおけるオープンエンド世代からオープンワールドインタラクションへ
- Abstract要約: オナイラ(英: Oneira)は、明示的な世界状態を通じて生成と相互作用の間のループを閉じるインタラクティブなビデオワールドモデルである。
オナイラは、長い地平線上での先行相互作用の効果を保ちながら、新たに生成されたオブジェクトとの直接的かつ一貫した相互作用を可能にする。
- 参考スコア(独自算出の注目度): 47.28440373515767
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: https://madaoer.github.io/projects/oneira
- Abstract(参考訳): 生成するビデオワールドモデルは、エージェントが簡単な方法でナビゲートし、対話できるオープンエンド環境を合成できるようになった。
しかし、オープンエンド世代は完全な相互作用を示さない: 生成された世界が拡大するにつれて、ナビゲーションによって新しく作成されたコンテンツはエージェントができることを拡張し、エージェントが世界を変えるにつれて、これらの変化は過渡的な視覚効果よりも環境の永続的な部分になる。
我々は,これら2つの要件を,新たに生成された,あるいは遭遇したエンティティを行動可能な世界に組み込むオープンワールド・インターアクティビティ(Open-World Interactive)と,相互作用の結果が世界状態にコミットされ,その後の観察や相互作用に影響を与え続ける永続状態(Persistent State)として特徴付ける。
我々は,コーディングエージェントが管理する明示的で拡張可能な世界状態を通じて,生成と相互作用のループを閉じるインタラクティブなビデオワールドモデルであるOneiraを紹介する。
エージェントは現在の観察と行動、あるいは高いレベルの目標を考慮し、世界状態を読み、関連するエンティティを根拠に、相互作用を計画し、その結果を世界状態表に書き戻す。
探索が新しい物体を明らかにすると、エージェントは生成された観測からそれらを取り込み、相互作用空間は生成された世界と膨張する。
一方、予め誘導された状態変化はビデオセグメント間で行われ、その後の世界進化における相互作用の持続的な部分の結果をもたらす。
更新された世界状態は、カメラ動作軌跡に沿って粗い条件付けビデオにレンダリングされ、映像生成装置は、その状態に表現されていない外観、動き、相互作用の詳細を埋める。
実験により、オナイラは新たに生成された物体との直接的かつ一貫した相互作用を可能にし、長い地平線上の先行相互作用の効果を保っていることが示された。
プロジェクトページ: https://madaoer.github.io/projects/oneira
関連論文リスト
- Code Plans, Diffusion Renders: Open-Ended Generative World Modeling [82.36042129562516]
我々は世界モデリングの新しいパラダイムである textbfCoDeR を紹介する。
本システムは,コードによる実行可能世界を明示的に構築し,映像生成モデルを用いて視覚的実現を行う。
論文 参考訳(メタデータ) (2026-09-22T14:12:15Z) - ActWorld: From Explorable to Interactive World Model via Action-Aware Memory [36.88820961480639]
本稿では,対話型世界モデルであるActWorldについて紹介する。
実験の結果、ActWorldは単一のモデル内でフレキシブルなナビゲーションとリッチなオブジェクトインタラクションの両方をサポートしています。
論文 参考訳(メタデータ) (2026-06-16T09:47:32Z) - Astra: General Interactive World Model with Autoregressive Denoising [73.6594791733982]
Astraはインタラクティブな汎用世界モデルであり、多様なシナリオのために現実世界の未来を生成する。
本稿では,自己回帰型認知型アーキテクチャを提案し,時間的因果的注意を用いて過去の観測を集約する。
Astraはインタラクティブで一貫性があり、一般的な長期的なビデオ予測を実現し、様々な形式のインタラクションをサポートする。
論文 参考訳(メタデータ) (2025-12-09T18:59:57Z) - InterDyn: Controllable Interactive Dynamics with Video Diffusion Models [50.38647583839384]
我々は、初期フレームと駆動対象またはアクターの動作を符号化する制御信号が与えられたインタラクティブな動画像を生成するフレームワークであるInterDynを提案する。
我々の重要な洞察は、大規模なビデオ生成モデルは、大規模ビデオデータからインタラクティブなダイナミクスを学習し、ニューラルと暗黙の物理シミュレーターの両方として機能できるということです。
論文 参考訳(メタデータ) (2024-12-16T13:57:02Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。