論文の概要: FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
- arxiv url: http://arxiv.org/abs/2603.19598v1
- Date: Fri, 20 Mar 2026 03:15:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-23 19:48:38.962121
- Title: FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
- Title(参考訳): FlowScene:マルチモーダルグラフ整流によるスタイル一貫性のある室内シーン生成
- Authors: Zhifei Yang, Guangyao Zhai, Keyang Lu, YuYang Yin, Chao Zhang, Zhen Xiao, Jieyi Long, Nassir Navab, Yikai Wang,
- Abstract要約: FlowSceneはグラフベースのシーン生成モデルで、シーンレイアウト、オブジェクト形状、オブジェクトテクスチャを協調的に生成する。
コアには、生成時にオブジェクト情報を交換する密結合整流モデルがあり、グラフ全体の協調推論を可能にする。
大規模な実験により、FlowSceneは、生成リアリズム、スタイルの整合性、人間の好みとの整合性という点で、言語条件とグラフ条件ベースラインの両方を上回ります。
- 参考スコア(独自算出の注目度): 52.01563659850181
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences.
- Abstract(参考訳): シーン生成は、高度な現実主義と、幾何学と外観の精密な制御の両方を要求する広範な産業的応用を持っている。
言語駆動の検索手法は、大きなオブジェクトデータベースからもっともらしいシーンを構成するが、オブジェクトレベルの制御を見落とし、しばしばシーンレベルの一貫性を強制できない。
グラフベースの定式化は、オブジェクトに対して高い制御性を提供し、関係を明示的にモデル化することで全体論的整合性を知らせるが、既存の手法は高忠実なテクスチャ化結果の生成に苦慮しているため、実用性が制限される。
本研究では,シーンレイアウト,オブジェクト形状,オブジェクトテクスチャを協調的に生成するマルチモーダルグラフ上に条件付き三分岐シーン生成モデルであるFlowSceneを提案する。
コアには、生成時にオブジェクト情報を交換する密結合整流モデルがあり、グラフ全体の協調推論を可能にする。
これにより、オブジェクトの形状、テクスチャ、関係を微妙に制御し、シーンレベルのスタイルコヒーレンスを構造や外観にわたって強化することができる。
大規模な実験により、FlowSceneは、生成リアリズム、スタイルの整合性、人間の好みとの整合性という点で、言語条件とグラフ条件ベースラインの両方を上回ります。
関連論文リスト
- CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph
Diffusion [83.30168660888913]
シーングラフを対応する制御可能な3Dシーンに変換する完全生成モデルであるCommonScenesを提案する。
パイプラインは2つのブランチで構成されており、1つは変分オートエンコーダでシーン全体のレイアウトを予測し、もう1つは互換性のある形状を生成する。
生成されたシーンは、入力シーングラフを編集し、拡散モデルのノイズをサンプリングすることで操作することができる。
論文 参考訳(メタデータ) (2023-05-25T17:39:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。