論文の概要: MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
- arxiv url: http://arxiv.org/abs/2604.08364v1
- Date: Thu, 09 Apr 2026 15:29:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-10 18:34:05.995129
- Title: MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
- Title(参考訳): MegaStyle: 一貫性のあるテキスト・画像間のスタイルマッピングによる多元的およびスケーラブルなスタイルデータセットの構築
- Authors: Junyao Gao, Sibo Liu, Jiaxing Li, Yanan Sun, Yuanpeng Tu, Fei Shen, Weidong Zhang, Cairong Zhao, Jun Zhang,
- Abstract要約: 私たちは、新しいスケーラブルなデータキュレーションパイプラインであるMegaStyleを紹介します。
我々は170Kスタイルのプロンプトと400Kコンテンツプロンプトを備えた多種多様なバランスの取れたプロンプトギャラリーをキュレートし、大規模スタイルのデータセットMegaStyle-1.4Mを生成する。
実験は、スタイルデータセットにおけるスタイル内の一貫性、スタイル間の多様性、高品質を維持することの重要性と、提案したMegaStyle-1.4Mの有効性を示す。
- 参考スコア(独自算出の注目度): 42.01890884376506
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and high-quality style dataset. We achieve this by leveraging the consistent text-to-image style mapping capability of current large generative models, which can generate images in the same style from a given style description. Building on this foundation, we curate a diverse and balanced prompt gallery with 170K style prompts and 400K content prompts, and generate a large-scale style dataset MegaStyle-1.4M via content-style prompt combinations. With MegaStyle-1.4M, we propose style-supervised contrastive learning to fine-tune a style encoder MegaStyle-Encoder for extracting expressive, style-specific representations, and we also train a FLUX-based style transfer model MegaStyle-FLUX. Extensive experiments demonstrate the importance of maintaining intra-style consistency, inter-style diversity and high-quality for style dataset, as well as the effectiveness of the proposed MegaStyle-1.4M. Moreover, when trained on MegaStyle-1.4M, MegaStyle-Encoder and MegaStyle-FLUX provide reliable style similarity measurement and generalizable style transfer, making a significant contribution to the style transfer community. More results are available at our project website https://jeoyal.github.io/MegaStyle/.
- Abstract(参考訳): 本稿では,新しい,スケーラブルなデータキュレーションパイプラインであるMegaStyleを紹介する。
これを実現するために,既存の大規模生成モデルの一貫したテキスト・ツー・イメージ・スタイルマッピング機能を活用し,与えられたスタイル記述から同じスタイルの画像を生成する。
この基盤を基盤として、170Kスタイルのプロンプトと400Kコンテンツプロンプトを備えた多種多様なバランスの取れたプロンプトギャラリーをキュレートし、コンテンツスタイルのプロンプトの組み合わせによって大規模スタイルのデータセットMegaStyle-1.4Mを生成する。
MegaStyle-1.4Mでは,表現型,スタイル固有の表現を抽出するためのスタイル教師付きコントラスト学習を提案し,FLUXに基づくスタイルトランスファーモデルであるMegaStyle-FLUXを訓練する。
大規模な実験は、スタイルデータセットにおけるスタイル内の一貫性、スタイル間の多様性、高品質を維持することの重要性と、提案したMegaStyle-1.4Mの有効性を示す。
さらに,MegaStyle-1.4Mでトレーニングを行うと,MegaStyle-EncoderとMegaStyle-FLUXは信頼性の高いスタイル類似度測定と一般化可能なスタイル転送を提供し,スタイル転送コミュニティに大きな貢献をする。
さらなる結果は、プロジェクトのWebサイトhttps://jeoyal.github.io/MegaStyle/.comで公開されています。
関連論文リスト
- TeleStyle: Content-Preserving Style Transfer in Images and Videos [52.76027947278353]
画像とビデオの両方をスタイリングするための軽量モデルであるTeleStyleを提示する。
異なるスタイルの高品質なデータセットをキュレートし、数千の多様性のあるイン・ザ・ワイルドなスタイルのカテゴリを使用してトリプレットを合成した。
TeleStyleは、スタイルの類似性、コンテントの一貫性、美的品質という、3つの中核評価指標で最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2026-01-28T02:16:03Z) - OmniStyle: Filtering High Quality Style Transfer Data at Scale [22.88223293456666]
OmniStyle-1Mは,100万以上のコンテンツスタイルスティル化画像三重項からなる大規模ペア型転送データセットである。
我々は,OmniStyle-1Mが教師付きトレーニングを通じて,効率よくスケーラブルなスタイル転送モデルを実現するだけでなく,ターゲットのスタイリゼーションを正確に制御できることを示す。
論文 参考訳(メタデータ) (2025-05-20T07:29:21Z) - Pluggable Style Representation Learning for Multi-Style Transfer [41.09041735653436]
スタイルモデリングと転送を分離してスタイル転送フレームワークを開発する。
スタイルモデリングでは,スタイル情報をコンパクトな表現に符号化するスタイル表現学習方式を提案する。
スタイル転送のために,プラガブルなスタイル表現を用いて多様なスタイルに適応するスタイル認識型マルチスタイル転送ネットワーク(SaMST)を開発した。
論文 参考訳(メタデータ) (2025-03-26T09:44:40Z) - StyleShot: A Snapshot on Any Style [20.41380860802149]
テスト時間チューニングを伴わない汎用的なスタイル転送には,優れたスタイル表現が不可欠であることを示す。
スタイル認識型エンコーダと、StyleGalleryと呼ばれるよく編成されたスタイルデータセットを構築することで、これを実現する。
当社のアプローチであるStyleShotは,テストタイムチューニングを必要とせずに,さまざまなスタイルを模倣する上で,シンプルかつ効果的なものです。
論文 参考訳(メタデータ) (2024-07-01T16:05:18Z) - StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter [78.75422651890776]
StyleCrafterは、トレーニング済みのT2Vモデルをスタイルコントロールアダプタで拡張する汎用的な方法である。
コンテンツスタイルのゆがみを促進するため,テキストプロンプトからスタイル記述を取り除き,参照画像のみからスタイル情報を抽出する。
StyleCrafterは、テキストの内容と一致し、参照画像のスタイルに似た高品質なスタイリングビデオを効率よく生成する。
論文 参考訳(メタデータ) (2023-12-01T03:53:21Z) - Domain Enhanced Arbitrary Image Style Transfer via Contrastive Learning [84.8813842101747]
Contrastive Arbitrary Style Transfer (CAST) は、新しいスタイル表現学習法である。
本フレームワークは,スタイルコード符号化のための多層スタイルプロジェクタ,スタイル分布を効果的に学習するためのドメイン拡張モジュール,画像スタイル転送のための生成ネットワークという,3つのキーコンポーネントから構成される。
論文 参考訳(メタデータ) (2022-05-19T13:11:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。