論文の概要: CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets
- arxiv url: http://arxiv.org/abs/2608.14403v1
- Date: Fri, 14 Aug 2026 15:44:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-17 20:14:29.425193
- Title: CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets
- Title(参考訳): CRAFT:構成対象を含まない主観的パーソナライゼーションのための注意細工による制約付きリワード
- Abstract要約: Intention Fine-Tuningによる制約付きリワード(Constrained Reward via Attention Fine-Tuning)は,LoRAアダプタを介して事前学習した参照認識MMDiTを微調整する単一ステップのReFLフレームワークである。
CRAFTはXBench rev上での最先端のパフォーマンスを達成し、コンポジションターゲットの監視は行わない。
- 参考スコア(独自算出の注目度): 21.908108950896334
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene. Producing such targets demands a costly multi-stage curation pipeline---LLM-based prompt generation, T2I-based composed-target synthesis, reference-subject extraction, VLM-based quality filtering, and correspondence labeling---and tightly couples each method to a particular target synthesizer and curation choice. We introduce \emph{CRAFT} (Constrained Reward via Attention Fine-Tuning), a single-step ReFL framework that fine-tunes a pre-trained \emph{reference-aware} MMDiT via LoRA adapters using a compact reference-only data construction---$10$K reference images and subject masks, with no composed-target supervision. CRAFT realizes a \emph{Where to look} principle: attention-level rewards align noise- and phrase-token attention with the correct reference subject, and the resulting per-subject attention masks gate a pixel-level identity reward to keep image-space supervision consistent with the learned attention routing. Applied to FLUX.2-klein-9B, CRAFT achieves state-of-the-art performance on XVerseBench \rev{while using no composed-target supervision---only $10$K reference-only samples, whereas prior generalized methods require $150$K to over $2$M composed-target pairs}. The same recipe transfers to other reference-aware backbones, consistently improving performance. Project page: https://jihun999.github.io/projects/CRAFT/.
- Abstract(参考訳): テーマ駆動のイメージパーソナライゼーション-新しいシーンにおける1つまたは複数の参照対象のアイデンティティを保持する新しいイメージの生成-は、現代の視覚コンテンツ作成の基盤となる能力である。
現在は、数十万から数百万対のペア 'emph{(reference, composed-target") の例に対して、事前訓練されたマルチモーダル拡散変換器 (MMDiT) を微調整する一般化された手法によって支配されている。
このようなターゲットを作成するには、コストのかかる多段階のキュレーションパイプライン--LLMベースのプロンプト生成、T2Iベースのコンポジションターゲット合成、参照オブジェクト抽出、VLMベースの品質フィルタリング、対応ラベル作成が必要です。
Intention Fine-Tuning, a single-step ReFL framework, a single-step ReFL framework that a pre-trained \emph{reference-aware} MMDiT via LoRA adapters using a compact reference-only data construction--10$K Reference image and subject masks with no composed-target supervision。
CRAFTは、注目レベルの報酬は、ノイズやフレーズを意識した注意と正しい参照対象とを一致させ、その結果、オブジェクトごとの注目マスクは、学習された注目経路と整合した画像空間の監督を維持するために、ピクセルレベルのアイデンティティ報酬をゲートする、という、emph{Where to look}原則を実現している。
FLUX.2-klein-9B に適用された CRAFT は、XVerseBench \rev{ While の最先端のパフォーマンスを、コンストラクトターゲットの監視無しで達成する。
同じレシピが他の参照対応バックボーンに転送され、継続的にパフォーマンスが向上する。
プロジェクトページ:https://jihun999.github.io/projects/CRAFT/。
関連論文リスト
- UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation [65.53694602893042]
VLMエンコーディングの前にVTとVAE機能を融合した統合ビジュアルコンディショニングフレームワークを提案する。
2つのマルチ参照生成ベンチマークの実験により、UniCustomは主題の一貫性、命令従順、構成の忠実さを一貫して改善することを示した。
論文 参考訳(メタデータ) (2026-05-12T13:10:05Z) - MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval [32.33545237942899]
Composed Image Retrieval (CIR) は、ギャラリーから、参照画像と修正テキストを使用してターゲット画像を取得するタスクである。
トレーニング不要なゼロショットCIRフレームワークとして再ランク付けされたChain-of-Thought(MCoT-RE)を提案する。
論文 参考訳(メタデータ) (2025-07-17T06:22:49Z) - Halton Scheduler For Masked Generative Image Transformer [51.82285573627426]
Masked Generative Image Transformers (MaskGIT)はスケーラブルで効率的な画像生成フレームワークとして登場した。
トークン間の相互情報に基づいて,MaskGITにおけるサンプリング対象を解析する。
そこで本研究では,最初の信頼性スケジューラの代わりに,Haltonスケジューラに基づく新しいサンプリング戦略を提案する。
論文 参考訳(メタデータ) (2025-03-21T12:00:59Z) - Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning [18.217337720633076]
我々は、信頼できる報酬信号を提供する$lambda$-Harmonic reward関数を提示する。
提案アルゴリズムは,最新のCLIP-Iスコア0.833,CLIP-Tスコア0.314をDreamBench上で達成する。
論文 参考訳(メタデータ) (2024-07-16T20:40:25Z) - A Fair Ranking and New Model for Panoptic Scene Graph Generation [51.78798765130832]
Decoupled SceneFormer(DSFormer)は、既存のすべてのシーングラフモデルよりも優れた2段階モデルである。
基本設計原則として、DSFormerは被写体とオブジェクトマスクを直接特徴空間にエンコードする。
論文 参考訳(メタデータ) (2024-07-12T12:28:08Z) - SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation [11.243400478302771]
Referring Expression Consistency (RES) は、テキストによって参照される画像において、対象オブジェクトのセグメンテーションマスクを提供することを目的としている。
アルゴリズムの革新を取り入れたRESのための弱教師付きブートストラップアーキテクチャを提案する。
論文 参考訳(メタデータ) (2024-07-02T16:02:25Z) - Multi-modal Generation via Cross-Modal In-Context Learning [50.45304937804883]
複雑なマルチモーダルプロンプトシーケンスから新しい画像を生成するMGCC法を提案する。
我々のMGCCは、新しい画像生成、マルチモーダル対話の促進、テキスト生成など、多種多様なマルチモーダル機能を示している。
論文 参考訳(メタデータ) (2024-05-28T15:58:31Z) - LeftRefill: Filling Right Canvas based on Left Reference through
Generalized Text-to-Image Diffusion Model [55.20469538848806]
leftRefillは、参照誘導画像合成のための大規模なテキスト・ツー・イメージ(T2I)拡散モデルを利用する革新的なアプローチである。
本稿では、参照誘導画像合成に大規模なテキスト・ツー・イメージ拡散モデル(T2I)を効果的に活用するための革新的なアプローチであるLeftRefillを紹介する。
論文 参考訳(メタデータ) (2023-05-19T10:29:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。