論文の概要: PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation
- arxiv url: http://arxiv.org/abs/2604.03611v1
- Date: Sat, 04 Apr 2026 06:50:51 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-07 15:49:18.671198
- Title: PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation
- Title(参考訳): PortraitCraft: ポートレートコンポジションの理解と生成のためのベンチマーク
- Authors: Yuyang Sha, Zijie Lou, Youyun Tang, Xiaochao Qu, Haoxiang Li, Ting Liu, Luoqi Liu,
- Abstract要約: PortraitCraftは、ポートレートコンポジションの理解と生成のための統一されたベンチマークである。
PortraitCraftは、約5万枚の実際のポートレート画像のデータセット上に構築されている。
- 参考スコア(独自算出の注目度): 21.26852995282223
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This limits systematic research on structured portrait composition analysis and controllable portrait generation under explicit composition requirements. In this paper, we introduce PortraitCraft, a unified benchmark for portrait composition understanding and generation. PortraitCraft is built on a dataset of approximately 50,000 curated real portrait images with structured multi-level supervision, including global composition scores, annotations over 13 composition attributes, attribute-level explanation texts, visual question answering pairs, and composition-oriented textual descriptions for generation. Based on this dataset, we establish two complementary benchmark tasks for composition understanding and composition-aware generation within a unified framework. The first evaluates portrait composition understanding through score prediction, fine-grained attribute reasoning, and image-grounded visual question answering, while the second evaluates portrait generation from structured composition descriptions under explicit composition constraints. We further define standardized evaluation protocols and provide reference baseline results with representative multimodal models. PortraitCraft provides a comprehensive benchmark for future research on fine-grained portrait understanding, interpretable aesthetic assessment, and controllable portrait generation.
- Abstract(参考訳): ポートレート構成は、肖像画の美学と視覚コミュニケーションにおいて中心的な役割を果たすが、既存のデータセットとベンチマークは主に粗い美的評価、一般的な画像の美学、あるいは制約のない肖像画生成に焦点を当てている。
この制限は、明示的な組成要求下での構造的肖像画合成分析と制御可能な肖像画生成に関する体系的研究である。
本稿では,ポートレートクラフト(PortraitCraft)について紹介する。
PortraitCraftは、グローバルなコンポジションスコア、13のコンポジション属性上のアノテーション、属性レベルの説明テキスト、視覚的な質問応答ペア、生成のためのコンポジション指向のテキスト記述などを含む、構造化されたマルチレベル監視を備えた、約50,000のキュレートされた実際のポートレートイメージのデータセット上に構築されている。
このデータセットに基づいて、構成理解と構成認識のための2つの補完的なベンチマークタスクを統一されたフレームワーク内に構築する。
第1は、スコア予測、微粒化属性推論、画像地上視覚質問応答による肖像画構成理解、第2は、明示的な構成制約の下で構成された構成記述からの肖像画生成を評価する。
さらに、標準化された評価プロトコルを定義し、代表的マルチモーダルモデルによる基準ベースライン結果を提供する。
PortraitCraftは、詳細なポートレート理解、解釈可能な美的評価、制御可能なポートレート生成に関する将来の研究のための総合的なベンチマークを提供する。
関連論文リスト
- A Sketch+Text Composed Image Retrieval Dataset for Thangka [14.600552992453977]
Composed Image Retrieval (CIR)は、複数のクエリーモダリティを組み合わせることで画像検索を可能にする。
CIRThanは、Thangkaイメージ用のスケッチ+テキストコンポジションイメージ検索データセットである。
論文 参考訳(メタデータ) (2026-02-09T09:14:29Z) - Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception [101.76154325436544]
マルチモーダルな大規模言語モデル (MLLM) は、既存の低レベルビジョンベンチマークで顕著な性能を示している。
Q-Bench-Portraitは、画像品質の知覚に特化して設計された最初の総合的なベンチマークである。
論文 参考訳(メタデータ) (2026-01-26T10:37:20Z) - Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition [61.956088652094515]
視覚と言語モデル(VLM)は、驚くべきゼロショット認識能力を示した。
しかし、それらは視覚言語的構成性、特に言語的理解ときめ細かい画像テキストアライメントの課題に直面している。
本稿では,構成性と認識の複雑な関係について考察する。
論文 参考訳(メタデータ) (2024-06-13T17:58:39Z) - T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation [55.16845189272573]
T2I-CompBench++は、合成テキスト・画像生成のための拡張ベンチマークである。
8000のコンポジションテキストプロンプトは、属性バインディング、オブジェクト関係、生成数、複雑なコンポジションの4つのグループに分類される。
論文 参考訳(メタデータ) (2023-07-12T17:59:42Z) - Portrait Interpretation and a Benchmark [49.484161789329804]
提案した肖像画解釈は,人間の知覚を新たな体系的視点から認識する。
我々は,身元,性別,年齢,体格,身長,表情,姿勢をラベル付けした25万枚の画像を含む新しいデータセットを構築した。
筆者らの実験結果から, 肖像画解釈に関わるタスクを組み合わせることで, メリットが得られることが示された。
論文 参考訳(メタデータ) (2022-07-27T06:25:09Z) - StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis [52.341186561026724]
構成性の欠如は、堅牢性と公正性に深刻な影響を及ぼす可能性がある。
テキスト対画像合成の合成性を改善するための新しいフレームワークであるStyleT2Iを導入する。
その結果,StyleT2Iは入力テキストと合成画像との整合性という点で従来の手法よりも優れていた。
論文 参考訳(メタデータ) (2022-03-29T17:59:50Z) - NPRportrait 1.0: A Three-Level Benchmark for Non-Photorealistic
Rendering of Portraits [67.58044348082944]
本稿では,スタイリングされたポートレート画像の評価のための,新しい3レベルベンチマークデータセットを提案する。
厳密な基準が構築に使われ、その一貫性はユーザスタディによって検証された。
ポートレート・スタイル化アルゴリズムを評価するための新しい手法が開発されている。
論文 参考訳(メタデータ) (2020-09-01T18:04:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。