論文の概要: TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation
- arxiv url: http://arxiv.org/abs/2606.20709v1
- Date: Tue, 16 Jun 2026 10:45:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-26 15:52:30.095159
- Title: TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation
- Title(参考訳): テレスタイルV2: 自己蒸留と分散マッチ蒸留によるコンテンツ保存型トランスファーを超えて
- Authors: Shiwen Zhang, Yifan Xu, Haibin Huang, Chi Zhang, Xuelong Li,
- Abstract要約: コンテンツ参照とスタイル参照が与えられた場合、コンテンツ保存スタイル転送は、スタイル化された出力を生成するモデルを必要とする。
TeleStyle V1は、フォトリアリスティックなコンテンツ参照と芸術スタイル参照で訓練されている。
TeleStyle V2は、Realistic-and-Realistic(RnR)、Realistic-and-Stylized(RnS)、Stylized-and-Stylized(SnS)の形式でContent-Style参照をサポートする。
- 参考スコア(独自算出の注目度): 54.141830240935455
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Given a content reference and a style reference, content-preserving style transfer requires the model to generate stylized outputs with content and style consistency. We introduced TeleStyle V1 to tackle this problem. However, TeleStyle V1 is trained with photorealistic content reference and artistic style reference, which makes it incapable to cope with artistic content reference and realistic style reference in most cases. In this paper, we designed a Self-Distillation data synthesis strategy to construct such triplets from TeleStyle V1. Trained with such self-distilled triplets, our TeleStyle V2 supports Content-Style references in the forms of Realistic-and-Realistic (RnR), Realistic-and-Stylized (RnS), Stylized-and-Realistic (SnR), Stylized-and-Stylized (SnS). In addition, we found Distribution Matching Distillation could preserve the general text-guided image editing capability of the foundation model and fix the content consistency degradation caused by SFT process. Through quantitative evaluations, our TeleStyleV2-QIE-2509-DMD performs at least on par with Qwen-Image-Edit-2509-DMD, demonstrating strong general image editing skills beyond content-preserving style transfer. We observed the content/style reference order confusion problem in TeleStyle V1 and further introduced prompt enhancer to solve it. TeleStyle V2 uses Qwen-Image-Edit's VLM encoder, Qwen2.5-VL-7B, to generate content prompt and style prompt for free. TeleStyle V2 could achieve comparable style transfer performance with state-of-the-art commercial model, gemini-3-pro-image-preview.
- Abstract(参考訳): コンテンツ参照とスタイル参照が与えられた場合、コンテンツ保存スタイル転送は、コンテンツとスタイル整合性を備えたスタイル化されたアウトプットを生成するモデルを必要とする。
私たちはこの問題に対処するためにTeleStyle V1を導入しました。
しかし、TeleStyle V1は、フォトリアリスティックなコンテンツ参照と芸術的なスタイル参照で訓練されており、ほとんどの場合、芸術的なコンテンツ参照と現実的なスタイル参照に対処することができない。
本稿では,TeleStyle V1からのトリプレット構築のための自己蒸留データ合成戦略を考案した。
我々のTeleStyle V2は、自己蒸留三重項を用いて訓練され、Realistic-and-Realistic(RnR)、Realistic-and-Stylized(RnS)、Stylized-and-Realistic(SnR)、Stylized-and-Stylized(SnS)の形式でContent-Style参照をサポートします。
さらに, 分散マッチング蒸留は, 基礎モデルの一般的なテキスト誘導画像編集能力を保ち, SFTプロセスによるコンテンツ一貫性の低下を解消できることがわかった。
定量的評価により、我々のTeleStyleV2-QIE-2509-DMDは少なくともQwen-Image-Edit-2509-DMDと同等に動作し、コンテンツ保存スタイル転送を超える強力な画像編集技術を示している。
我々は、TeleStyle V1におけるコンテンツ/スタイル参照順序混乱問題を観察し、さらにプロンプトエンハンサーを導入して解決した。
TeleStyle V2はQwen-Image-EditのVLMエンコーダであるQwen2.5-VL-7Bを使用して、コンテンツプロンプトとスタイルプロンプトを無償で生成する。
TeleStyle V2は最先端の商用モデルである gemini-3-pro-image-preview で同等のスタイルの転送性能を達成できた。
関連論文リスト
- TeleStyle: Content-Preserving Style Transfer in Images and Videos [52.76027947278353]
画像とビデオの両方をスタイリングするための軽量モデルであるTeleStyleを提示する。
異なるスタイルの高品質なデータセットをキュレートし、数千の多様性のあるイン・ザ・ワイルドなスタイルのカテゴリを使用してトリプレットを合成した。
TeleStyleは、スタイルの類似性、コンテントの一貫性、美的品質という、3つの中核評価指標で最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2026-01-28T02:16:03Z) - QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit [54.11909509184315]
本稿では,Qwen-Image-Editでトレーニングされた最初のコンテンツ保存スタイル転送モデルを提案する。
QwenStyle V1は、スタイルの類似性、コンテントの一貫性、美的品質の3つのコアメトリクスで、最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2026-01-08T10:22:51Z) - DreamStyle: A Unified Framework for Video Stylization [18.820518165759403]
ビデオスタイリングのための統合フレームワークDreamStyleを紹介する。
1)テキスト誘導、(2)スタイル誘導、(3)ファーストフレーム誘導ビデオスタイリングをサポートする。
質的および定量的な評価は、DreamStyleが3つのビデオスタイリングタスク全てに適していることを示している。
論文 参考訳(メタデータ) (2026-01-06T07:42:12Z) - StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter [78.75422651890776]
StyleCrafterは、トレーニング済みのT2Vモデルをスタイルコントロールアダプタで拡張する汎用的な方法である。
コンテンツスタイルのゆがみを促進するため,テキストプロンプトからスタイル記述を取り除き,参照画像のみからスタイル情報を抽出する。
StyleCrafterは、テキストの内容と一致し、参照画像のスタイルに似た高品質なスタイリングビデオを効率よく生成する。
論文 参考訳(メタデータ) (2023-12-01T03:53:21Z) - StyleAdapter: A Unified Stylized Image Generation Model [97.24936247688824]
StyleAdapterは、様々なスタイリング画像を生成することができる統一型スタイリング画像生成モデルである。
T2I-adapter や ControlNet のような既存の制御可能な合成手法と統合することができる。
論文 参考訳(メタデータ) (2023-09-04T19:16:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。