論文の概要: StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
- arxiv url: http://arxiv.org/abs/2604.12575v1
- Date: Tue, 14 Apr 2026 10:55:43 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-15 19:11:32.401946
- Title: StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
- Title(参考訳): StructDiff:単一画像生成のための構造保存・空間制御可能な拡散モデル
- Authors: Yinxi He, Kang Liao, Chunyu Lin, Tianyi Wei, Yao Zhao,
- Abstract要約: StructDiffは、単一画像生成のための単一スケール拡散モデルに基づく生成フレームワークである。
3次元位置符号化(PE)を空間的先行として組み込んでおり、生成されたオブジェクトの位置、スケール、局所的な詳細を柔軟に制御することができる。
また、テキスト誘導画像生成、画像編集、アウトペインティング、ペイント・ツー・イメージ合成など、下流タスクにも幅広い適用性を示す。
- 参考スコア(独自算出の注目度): 72.84181869780627
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation aims to synthesize diverse samples with similar visual content to the source image by capturing its internal statistics, without relying on external data. However, existing methods often struggle to preserve the structural layout, especially for images with large rigid objects or strict spatial constraints. Moreover, most approaches lack spatial controllability, making it difficult to guide the structure or placement of generated content. To address these challenges, StructDiff introduces an \textit{adaptive receptive field} module to maintain both global and local distributions. Building on this foundation, StructDiff incorporates 3D positional encoding (PE) as a spatial prior, allowing flexible control over positions, scale, and local details of generated objects. To our knowledge, this spatial control capability represents the first exploration of PE-based manipulation in single-image generation. Furthermore, we propose a novel evaluation criterion for single-image generation based on large language models (LLMs). This criterion specifically addresses the limitations of existing objective metrics and the high labor costs associated with user studies. StructDiff also demonstrates broad applicability across downstream tasks, such as text-guided image generation, image editing, outpainting, and paint-to-image synthesis. Extensive experiments demonstrate that StructDiff outperforms existing methods in structural consistency, visual quality, and spatial controllability. The project page is available at https://butter-crab.github.io/StructDiff/.
- Abstract(参考訳): 本稿では,単一画像生成のための単一スケール拡散モデルに基づく生成フレームワークであるStructDiffを紹介する。
単一画像生成は、外部データに頼ることなく、内部統計をキャプチャすることで、ソース画像と類似した視覚的内容の多様なサンプルを合成することを目的としている。
しかし、既存の手法は、特に大きな剛体物体や厳密な空間的制約のある画像の場合、構造的レイアウトを維持するのに苦労することが多い。
さらに、ほとんどのアプローチは空間制御性に欠けており、生成されたコンテンツの構造や配置を導くことは困難である。
これらの課題に対処するため、StructDiffは、グローバルとローカルの両方のディストリビューションを維持するために、 \textit{adaptive Receptive Field}モジュールを導入した。
この基盤の上に、StructDiffは3D位置符号化(PE)を空間的先行として組み込んでおり、生成されたオブジェクトの位置、スケール、局所的な詳細を柔軟に制御することができる。
我々の知る限り、この空間制御能力は、単一画像生成におけるPEベースの操作の最初の探索である。
さらに,大規模言語モデル(LLM)に基づく単一画像生成のための新しい評価基準を提案する。
この基準は、既存の客観的指標の限界と、ユーザスタディに関連する高い労働コストに特に対処する。
StructDiffはまた、テキスト誘導画像生成、画像編集、アウトペイント、ペイント・ツー・イメージ合成など、下流タスクにまたがる幅広い適用性を示している。
広範囲な実験により、StructDiffは、構造的整合性、視覚的品質、空間的可制御性において、既存の手法よりも優れていることが示された。
プロジェクトのページはhttps://butter-crab.github.io/StructDiff/.comで公開されている。
関連論文リスト
- HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation [11.087309945227826]
レイアウト・画像生成のためのtextbfHierarchical textbfControllable (HiCo) 拡散モデルを提案する。
我々の重要な洞察は、レイアウトの階層的モデリングを通じて空間的ゆがみを実現することである。
自然シーンにおける多目的制御可能なレイアウト生成の性能を評価するため,HiCo-7Kベンチマークを提案する。
論文 参考訳(メタデータ) (2024-10-18T09:36:10Z) - A Structure-Guided Diffusion Model for Large-Hole Image Completion [85.61681358977266]
画像中の大きな穴を埋める構造誘導拡散モデルを開発した。
本手法は,最先端の手法と比較して,優れた,あるいは同等の視覚的品質を実現する。
論文 参考訳(メタデータ) (2022-11-18T18:59:01Z) - Image Inpainting via Conditional Texture and Structure Dual Generation [26.97159780261334]
本稿では, 構造制約によるテクスチャ合成とテクスチャ誘導による構造再構築をモデル化した, 画像インペイントのための新しい2ストリームネットワークを提案する。
グローバルな一貫性を高めるため、双方向Gated Feature Fusion (Bi-GFF)モジュールは構造情報とテクスチャ情報を交換・結合するように設計されている。
CelebA、Paris StreetView、Places2データセットの実験は、提案手法の優位性を実証している。
論文 参考訳(メタデータ) (2021-08-22T15:44:37Z) - Generating Diverse Structure for Image Inpainting With Hierarchical
VQ-VAE [74.29384873537587]
本稿では,異なる構造を持つ複数の粗い結果を第1段階で生成し,第2段階ではテクスチャを増補して各粗い結果を別々に洗練する,多彩な塗布用2段階モデルを提案する。
CelebA-HQ, Places2, ImageNetデータセットによる実験結果から,本手法は塗布ソリューションの多様性を向上するだけでなく,生成した複数の画像の視覚的品質も向上することが示された。
論文 参考訳(メタデータ) (2021-03-18T05:10:49Z) - MOGAN: Morphologic-structure-aware Generative Learning from a Single
Image [59.59698650663925]
近年,1つの画像のみに基づく生成モデルによる完全学習が提案されている。
多様な外観のランダムなサンプルを生成するMOGANというMOrphologic-structure-aware Generative Adversarial Networkを紹介します。
合理的な構造の維持や外観の変化など、内部機能に重点を置いています。
論文 参考訳(メタデータ) (2021-03-04T12:45:23Z) - Person-in-Context Synthesiswith Compositional Structural Space [59.129960774988284]
本研究では,コンテキスト合成におけるtextbfPersons という新たな問題を提案する。
コンテキストは、形状情報を欠いたバウンディングボックスオブジェクトレイアウトで指定され、キーポイントによる人物のポーズは、わずかに注釈付けされている。
入力構造におけるスターク差に対処するため、各(コンテキスト/人物)入力を「共有構成構造空間」に意図的に合成する2つの別個の神経枝を提案した。
この構造空間は多レベル特徴変調戦略を用いて画像空間にデコードされ、自己学習される
論文 参考訳(メタデータ) (2020-08-28T14:33:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。