論文の概要: Face inpainting with Identity Preserving Latent Diffusion Models
- arxiv url: http://arxiv.org/abs/2605.16696v1
- Date: Fri, 15 May 2026 23:19:43 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-19 17:57:46.914128
- Title: Face inpainting with Identity Preserving Latent Diffusion Models
- Title(参考訳): 潜時拡散モデルを用いた顔の塗装
- Authors: João Santos, Carlos Santiago, Manuel Marques,
- Abstract要約: ID-ControlNetは、潜伏拡散モデルに基づいて構築された、ID保存のフェイスペイントフレームワークである。
我々は、生成した顔とターゲットの識別表現とのアライメントを明確に強制する、アイデンティティ一貫性と三重項損失トレーニング戦略を導入する。
- 参考スコア(独自算出の注目度): 10.348748950921667
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Face inpainting techniques recover missing or occluded facial regions in a visually realistic manner, but preserving the identity in the final output remains a fundamental challenge. Identity consistency is crucial for downstream applications such as face recognition, digital forensics, and human-computer interaction, where even subtle identity distortions can significantly degrade performance or trust. Although diffusion-based generative models have recently achieved remarkable progress in image inpainting, they often struggle to faithfully retain individual-specific facial characteristics. On the other hand, existing identity-aware methods typically rely on costly fine-tuning, auxiliary supervision, or exhibit limited robustness to diverse occlusions, poses, and facial variations. To address these limitations, we propose ID-ControlNet, an identity-preserving face inpainting framework built upon latent diffusion models. Based on ControlNet architecture, our approach conditions the diffusion process on facial identity embeddings extracted from a pretrained face recognition network. This design enables reconstruction of occluded facial regions while maintaining global facial coherence and identity fidelity. Furthermore, we introduce an identity consistency and triplet loss training strategy that explicitly enforces alignment between the generated face and the target identity representation. Extensive experiments on CelebA-HQ, FFHQ, and on a new E-Mask dataset demonstrate that ID-ControlNet significantly improves identity preservation over standard diffusion-based inpainting methods, achieving performance comparable to SOTA identity-aware approaches.
- Abstract(参考訳): 顔の塗装技術は、視覚的に現実的に顔領域の欠落や隠蔽を回復するが、最終的な出力における同一性を維持することは、依然として根本的な課題である。
アイデンティティの一貫性は、顔認識、デジタル法医学、人間とコンピュータの相互作用といった下流のアプリケーションにとって重要であり、微妙なアイデンティティの歪みでさえパフォーマンスや信頼を著しく低下させる可能性がある。
拡散に基づく生成モデルは近年、画像の塗布において顕著な進歩を遂げているが、しばしば個々の顔の特徴を忠実に保ち続けるのに苦労する。
一方、既存のアイデンティティ認識手法は、通常、コストのかかる微調整、補助的な監督、あるいは多様なオクルージョン、ポーズ、顔のバリエーションに対する限られた堅牢性を示す。
これらの制約に対処するため,潜時拡散モデル上に構築されたID保存顔インペイントフレームワークであるID-ControlNetを提案する。
ControlNetアーキテクチャに基づいて,事前学習した顔認識ネットワークから抽出した顔のアイデンティティ埋め込みの拡散過程を条件とした。
この設計は、グローバルな顔のコヒーレンスとアイデンティティの忠実さを維持しつつ、隠蔽された顔領域の再構築を可能にする。
さらに、生成した顔と対象の識別表現とのアライメントを明示的に強制する、アイデンティティ一貫性と三重項損失トレーニング戦略を導入する。
CelebA-HQ、FFHQ、および新しいE-Maskデータセットに対する大規模な実験により、ID-ControlNetは標準拡散法よりもID保存を著しく改善し、SOTAのアイデンティティ認識アプローチに匹敵するパフォーマンスを達成する。
関連論文リスト
- CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping [1.4323566945483497]
顔スワップは、ターゲット顔へのソース顔の同一性を活用することで、現実的な顔画像生成を最適化することを目的としている。
既存の方法、特にGANに基づく手法は、アイデンティティ保存と視覚リアリズムのバランスをとるのにしばしば苦労する。
本稿では、視線、アイデンティティ、顔解析を統合した最初の拡散型顔スワップアプローチであるCA-IDDを紹介する。
論文 参考訳(メタデータ) (2026-04-27T13:59:08Z) - SIDeR: Semantic Identity Decoupling for Unrestricted Face Privacy [53.75084833636302]
本稿では,非制限顔プライバシー保護のためのセマンティックデカップリング駆動フレームワークSIDeRを提案する。
SIDeRは、顔画像をマシン認識可能な識別特徴ベクトルと視覚的に知覚可能なセマンティックな外観成分に分解する。
認証されたアクセスのために、SIDeRは正しいパスワードが提供されるときに元の形式に復元できる。
論文 参考訳(メタデータ) (2026-02-04T19:30:48Z) - iFADIT: Invertible Face Anonymization via Disentangled Identity Transform [51.123936665445356]
顔の匿名化は、個人のプライバシーを保護するために顔の視覚的アイデンティティを隠すことを目的としている。
Invertible Face Anonymization の頭字語 iFADIT を Disentangled Identity Transform を用いて提案する。
論文 参考訳(メタデータ) (2025-01-08T10:08:09Z) - ID$^3$: Identity-Preserving-yet-Diversified Diffusion Models for Synthetic Face Recognition [60.15830516741776]
合成顔認識(SFR)は、実際の顔データの分布を模倣するデータセットを生成することを目的としている。
拡散燃料SFRモデルであるtextID3$を紹介します。
textID3$はID保存損失を利用して、多様だがアイデンティティに一貫性のある顔の外観を生成する。
論文 参考訳(メタデータ) (2024-09-26T06:46:40Z) - Disentangle Before Anonymize: A Two-stage Framework for Attribute-preserved and Occlusion-robust De-identification [55.741525129613535]
匿名化前の混乱」は、新しい二段階フレームワーク(DBAF)である
このフレームワークには、Contrastive Identity Disentanglement (CID)モジュールとKey-authorized Reversible Identity Anonymization (KRIA)モジュールが含まれている。
大規模な実験により,本手法は最先端の非識別手法より優れていることが示された。
論文 参考訳(メタデータ) (2023-11-15T08:59:02Z) - Controllable Inversion of Black-Box Face Recognition Models via
Diffusion [8.620807177029892]
我々は,事前学習した顔認識モデルの潜在空間を,完全なモデルアクセスなしで反転させる作業に取り組む。
本研究では,条件付き拡散モデル損失が自然発生し,逆分布から効果的にサンプル化できることを示す。
本手法は,生成過程を直感的に制御できる最初のブラックボックス顔認識モデル逆変換法である。
論文 参考訳(メタデータ) (2023-03-23T03:02:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。