論文の概要: Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System
- arxiv url: http://arxiv.org/abs/2609.04151v2
- Date: Thu, 10 Sep 2026 14:54:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-11 19:20:15.677341
- Title: Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System
- Title(参考訳): 生成画像モデルにおける永続的アイデンティティ保存:ベンチマークと評価システム
- Authors: Mengwei Ren, Xuaner Zhang, Zhihao Xia,
- Abstract要約: 生成画像モデルは高品質な画像を生成し、複雑な命令に従うことができ、正確な編集をサポートする。
しかし、誰が描かれているか、何が描かれているかを保存するのに苦戦している。
本研究は, アイデンティティの保存が, 現在の生成基盤モデルの明確な限界であることを示す。
- 参考スコア(独自算出の注目度): 20.281088110768955
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surrounding scene changes. Existing subject-driven methods make fundamentally different choices about where identity is represented: through the input context (GPT-Image-2, NB2), as trainable subject-specific model parameters (LoRA), or as a persistent identity layer (PHOTA IDENTITY) reusable across generations and edits. We systematically benchmark these paradigms across subject-driven generation, editing, restoration, and multi-subject settings, with tasks designed to increasingly stress identity preservation. Our results show that identity preservation remains a distinct limitation of current generative foundation models: strong image quality and instruction following do not necessarily imply strong identity fidelity, and identity degradation becomes more pronounced under iterative edits, small subject scales, severe image degradation, and multi-subject composition. Persistent identity substantially reduces this degradation across generation, editing, and restoration, consistently improving identity preservation when applied to different foundation models while maintaining comparable instruction adherence and perceptual image quality. These results suggest that identity does not simply emerge from increasingly capable generative models, but can instead be represented as persistent subject knowledge that is composed independently with the underlying generative model.
- Abstract(参考訳): 生成する画像モデルは、高品質な画像を生成し、複雑な指示に従い、正確な編集をサポートすることができるが、誰が何を描いたかを保存するのに苦戦している。
特定の被写体の画像を生成または編集する際には、ポーズ、表現、外観、視点、周囲のシーンの変化としてアイデンティティが漂うことがある。
既存の対象駆動方式では、入力コンテキスト(GPT-Image-2, NB2)、トレーニング可能な対象特化モデルパラメータ(LoRA)、世代間で再利用可能な永続ID層(PHOTA IDENTITY)など、アイデンティティの表現方法が根本的に異なる。
我々は、これらのパラダイムを主題駆動生成、編集、復元、多目的設定にわたって体系的にベンチマークし、アイデンティティの保存をますます強調するタスクを設計する。
画像の高画質化と指示が必ずしも強いアイデンティティの忠実さを示唆するものではないこと,また,反復的編集,小被写体スケール,重度画像劣化,多目的合成などによりアイデンティティの劣化が顕著になる。
永続的アイデンティティは、生成、編集、復元におけるこの劣化を著しく低減し、異なる基礎モデルに適用した際のアイデンティティ保存を一貫して改善し、同等の命令順守と知覚的画像品質を維持している。
これらの結果は、アイデンティティは、単に能力の高い生成モデルから生まれるのではなく、基礎となる生成モデルと独立して構成される永続的な主題的知識として表すことができることを示唆している。
関連論文リスト
- Latent-Identity Tuning in Text-to-Image Personalization Models [65.31471460269033]
テキスト・ツー・イメージのパーソナライズ・モデルにおける詳細なアイデンティティチューニング手法を提案する。
通常の画像編集とは異なり、アイデンティティチューニングは特定のアイデンティティの潜在表現を変更する。
この空間と選択されたトークンによって定義された部分空間内で有意な方向を識別できることが示される。
論文 参考訳(メタデータ) (2026-07-13T17:59:49Z) - WithAnyone: Towards Controllable and ID Consistent Image Generation [83.55786496542062]
アイデンティティ・一貫性・ジェネレーションは、テキスト・ツー・イメージ研究において重要な焦点となっている。
マルチパーソンシナリオに適した大規模ペアデータセットを開発する。
本稿では,データと多様性のバランスをとるためにペアデータを活用する,対照的なアイデンティティ損失を持つ新たなトレーニングパラダイムを提案する。
論文 参考訳(メタデータ) (2025-10-16T17:59:54Z) - EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance [20.430259028981094]
ゼロショットパーソナライズされた画像生成モデルは、与えられたテキストプロンプトと被写体画像の両方に一致した画像を作成することを目的としている。
既存の手法では、細かな被写体の詳細を捉えるのに苦労することが多く、一方のガイダンスを他方よりも優先することが多い。
EZIGenは、固定トレーニング済みのDiffusion UNet自体を主題エンコーダとして活用する。
論文 参考訳(メタデータ) (2024-09-12T14:44:45Z) - PortraitBooth: A Versatile Portrait Model for Fast Identity-preserved
Personalization [92.90392834835751]
PortraitBoothは高効率、堅牢なID保存、表現編集可能な画像生成のために設計されている。
PortraitBoothは計算オーバーヘッドを排除し、アイデンティティの歪みを軽減する。
生成した画像の多様な表情に対する感情認識のクロスアテンション制御が組み込まれている。
論文 参考訳(メタデータ) (2023-12-11T13:03:29Z) - DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven
Text-to-Image Generation [50.39533637201273]
主観駆動型テキスト・ツー・イメージ生成のためのID保存型アンタングル型チューニングフレームワークであるDisenBoothを提案する。
DisenBoothは、ID保存の埋め込みとアイデンティティ関連の埋め込みを組み合わせることで、より世代的柔軟性と制御性を示す。
論文 参考訳(メタデータ) (2023-05-05T09:08:25Z) - T-Person-GAN: Text-to-Person Image Generation with Identity-Consistency
and Manifold Mix-Up [16.165889084870116]
テキストのみに条件付けされた高解像度の人物画像を生成するためのエンドツーエンドアプローチを提案する。
2つの新しいメカニズムで人物画像を生成するための効果的な生成モデルを開発する。
論文 参考訳(メタデータ) (2022-08-18T07:41:02Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。