論文の概要: DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
- arxiv url: http://arxiv.org/abs/2607.02290v1
- Date: Thu, 02 Jul 2026 15:07:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.881016
- Title: DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
- Title(参考訳): DisciplineGen-1M:多分野のビジュアル生成と編集のための大規模データセット
- Authors: Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Yiguo He, Mohan Zhang, Leyao Gu, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Xuanhe Zhou, Zhihang Zhong, Junchi Yan, Xue Yang,
- Abstract要約: 本稿では,テキスト・ツー・イメージ生成と画像編集をサポートする100万規模のデータセットであるDisciplineGen-1Mを紹介する。
数学、物理学、化学、生物学、地理、コンピュータ科学、経済学、歴史、音楽、スポーツにまたがる1.2Mのサンプルを含んでいる。
DisciplineGen-1M をベースとして,テキスト・画像生成と画像編集の両面での規律インフォームド推論モデルを提案する。
- 参考スコア(独自算出の注目度): 65.66783293569955
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spatial relations. We introduce DisciplineGen-1M, a million-scale multidisciplinary dataset that supports text-to-image generation and image editing. It contains 1.2M samples spanning mathematics, physics, chemistry, biology, geography, computer science, economics, history, music, and sports. To construct the dataset, we design a scalable framework that combines vector-graphics rendering, OCR-based editing, curated programmatic synthesis, and large-scale text-to-image filtering. These pipelines produce captions, editing instructions, structured annotations, and paired images with controllable semantic differences. Building on DisciplineGen-1M, we further introduce a discipline-informed reasoning-generation model for both text-to-image generation and image editing. Experiments on discipline-related benchmarks, GenExam and GRADE, show substantial improvements over open-source baselines, while evaluations on general reasoning-informed benchmarks, WISE and RISE, further indicate broader transfer. The results suggest that large-scale structured academic visual data is a key ingredient for moving image generation from aesthetic plausibility toward verifiable knowledge-grounded visual creation. We will publicly release our dataset, model, and source code of the data curation pipeline to ensure reproducibility and benefit future research.
- Abstract(参考訳): 最近の画像生成・編集モデルは、視覚的に魅力的な自然画像を生成することができるが、対象画像が学際的な概念、象徴的構造、正確な空間関係に依存する知識集約図である場合、それらは信頼できないままである。
テキスト・ツー・イメージ生成と画像編集をサポートする百万規模のマルチディシプリナ・データセットであるDisciplineGen-1Mを紹介する。
数学、物理学、化学、生物学、地理、コンピュータ科学、経済学、歴史、音楽、スポーツにまたがる1.2Mのサンプルを含んでいる。
このデータセットを構築するために、ベクトルグラフレンダリング、OCRベースの編集、プログラム合成のキュレート、大規模テキスト・画像フィルタリングを組み合わせたスケーラブルなフレームワークを設計する。
これらのパイプラインは、キャプション、編集命令、構造化アノテーション、および制御可能なセマンティックな違いを持つペア画像を生成する。
また、DisciplineGen-1Mをベースとして、テキスト・画像生成と画像編集の両面での規律インフォームド推論モデルを導入する。
規律関連のベンチマークであるGenExamとGRADEの実験では、オープンソースベースラインよりも大幅に改善されている。
以上の結果から, 大規模構造化された学術的視覚データが, 審美的妥当性から, 検証可能な知識に基づく視覚生成へ画像生成を移行させる重要な要素であることが示唆された。
データキュレーションパイプラインのデータセット、モデル、ソースコードを公開して、再現性を確保し、将来の研究に役立てます。
関連論文リスト
- S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing [25.725824689793402]
本稿では,科学画像の理解,生成,編集のためのオープンウェイト統一マルチモーダルモデルを提案する。
S1-Omni-Imageは、バックボーンS1-VL-32Bの科学的なマルチモーダル推論に基づいている。
統一されたフレームワークで、科学的な画像理解、生成、編集をサポートする。
論文 参考訳(メタデータ) (2026-06-23T11:19:34Z) - Charts Are Not Images: On the Challenges of Scientific Chart Editing [66.38730113476677]
textitFigEditは、3万以上のサンプルからなる科学的フィギュア編集のベンチマークである。
私たちのベンチマークでは、ピクセルレベルの操作の重大な制限が示されています。
textitFigEdit をリリースすることにより,構造対応図形編集の体系的な進歩の実現を目指す。
論文 参考訳(メタデータ) (2025-11-30T06:13:48Z) - Factuality Matters: When Image Generation and Editing Meet Structured Visuals [46.627460447235855]
我々は、13万の高品質な構造化画像対からなる大規模データセットを構築した。
FLUX.1 KontextとVLMを統合する統一モデルを訓練する。
3段階のトレーニングカリキュラムは、プログレッシブな特徴アライメント、知識の注入、推論による生成を可能にする。
論文 参考訳(メタデータ) (2025-10-06T17:56:55Z) - Interleaving Reasoning for Better Text-to-Image Generation [83.69082794730664]
テキストベース思考と画像合成を交互に行うIRG(Interleaving Reasoning Generation)を提案する。
IRGを効果的に訓練するために,2つのサブゴールをターゲットにしたIRGL(Interleaving Reasoning Generation Learning)を提案する。
実験の結果、SoTAの性能はGenEval, WISE, TIIF, GenAI-Bench, OneIG-ENで5~10ポイント向上した。
論文 参考訳(メタデータ) (2025-09-08T17:56:23Z) - OCR-VQGAN: Taming Text-within-Image Generation [4.5718306968064635]
我々はOCR-VQGAN,画像エンコーダ,およびOCR事前学習機能を利用してテキスト知覚損失を最適化するデコーダを提案する。
我々は,OCR-VQGANの有効性を図形再構成の課題に関するいくつかの実験により実証した。
論文 参考訳(メタデータ) (2022-10-19T16:37:48Z) - Re-Imagen: Retrieval-Augmented Text-to-Image Generator [58.60472701831404]
検索用テキスト・ツー・イメージ・ジェネレータ(再画像)
検索用テキスト・ツー・イメージ・ジェネレータ(再画像)
論文 参考訳(メタデータ) (2022-09-29T00:57:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。