論文の概要: Reinforcing the Generation Order of Multimodal Masked Diffusion Models
- arxiv url: http://arxiv.org/abs/2607.08056v1
- Date: Thu, 09 Jul 2026 02:07:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-10 14:45:27.393418
- Title: Reinforcing the Generation Order of Multimodal Masked Diffusion Models
- Title(参考訳): マルチモーダルマスク付き拡散モデルの生成順序の強化
- Authors: Yidong Ouyang, Zhe Wang, Sourav Bhabesh, Dmitriy Bespalov,
- Abstract要約: グループ相対政策最適化(GRPO)を用いて学習可能な制御モジュールを導入し,生成順序を決定する。
この制御ブロックの学習は,拡散言語モデルにおけるテキスト・画像のアライメントとマルチモーダル理解の両方を大幅に改善することを示す。
- 参考スコア(独自算出の注目度): 6.5976113604335564
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications. In this work, we investigate the optimization of generation order for both text-to-image synthesis and multimodal understanding. We first establish that, unlike structured problems in language generation such as Sudoku puzzles, model logits alone are insufficient for determining optimal generation sequences in text-to-image generation and multimodal understanding. To address this challenge, we introduce a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order. Our results demonstrate that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs. In particular, it enhances the model's ability to capture fine-grained spatial relationships in generated images while also strengthening performance on multimodal reasoning and comprehension tasks. We evaluate our framework on GenEval, an object-focused benchmark for text-to-image alignment, where it achieves 4.08% relative improvements. In addition, experiments on VLMEvalKit confirm 4.85% relative improvements in multimodal understanding, highlighting the broad effectiveness of our approach.
- Abstract(参考訳): 拡散言語モデル(DLM)は近年,自然言語生成タスクにおいて大きな進歩を遂げている。
近年の研究では、適応トークン生成順序付けは、数学的推論およびコード合成アプリケーションの性能を著しく向上させることができることが示されている。
本研究では,テキストと画像の合成とマルチモーダル理解の両方において生成順序の最適化について検討する。
まず、スドゥークパズルのような言語生成の構造化問題とは異なり、モデルロジットだけではテキスト・画像生成やマルチモーダル理解において最適な生成シーケンスを決定するには不十分であることを示す。
この課題に対処するため,グループ相対政策最適化(GRPO)を用いて学習可能な制御モジュールを導入し,生成順序を決定する。
この制御ブロックの学習はDLMにおけるテキスト・画像のアライメントとマルチモーダル理解の両方を大幅に改善することを示す。
特に、モデルが生成した画像のきめ細かい空間関係を捉える能力を高め、マルチモーダル推論や理解タスクの性能を高める。
我々は、テキストと画像のアライメントのためのオブジェクト中心のベンチマークであるGenEvalのフレームワークを評価し、4.08%の相対的な改善を実現した。
さらに, VLMEvalKit実験により, マルチモーダル理解の相対的な改善が4.85%確認された。
関連論文リスト
- Growing Visual Generative Capacity for Pre-Trained MLLMs [60.826355079902505]
Bridgeは純粋な自己回帰統合MLLMであり、学習済みの視覚的理解モデルを生成能力で強化する。
本稿では,コンパクトなセマンティックトークンと微細なピクセルトークンを統合するセマンティック・ツー・ピクセルの離散表現を提案する。
論文 参考訳(メタデータ) (2025-10-02T00:40:02Z) - Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation [63.50827603618498]
マルチモーダル理解と生成のための統一型マスク付き拡散モデル(MDM)であるLavida-Oを提案する。
Lavida-Oは、画像レベルの理解、オブジェクトのグラウンド化、画像編集、高解像度のテキスト・ツー・イメージ合成を可能にする単一のフレームワークを提供する。
Lavida-Oは、RefCOCOオブジェクトグラウンド、GenEvalテキスト画像生成、ImgEdit画像編集など、幅広いベンチマークで最先端のパフォーマンスを実現している。
論文 参考訳(メタデータ) (2025-09-23T17:05:46Z) - Interleaving Reasoning for Better Text-to-Image Generation [83.69082794730664]
テキストベース思考と画像合成を交互に行うIRG(Interleaving Reasoning Generation)を提案する。
IRGを効果的に訓練するために,2つのサブゴールをターゲットにしたIRGL(Interleaving Reasoning Generation Learning)を提案する。
実験の結果、SoTAの性能はGenEval, WISE, TIIF, GenAI-Bench, OneIG-ENで5~10ポイント向上した。
論文 参考訳(メタデータ) (2025-09-08T17:56:23Z) - SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards [55.99492656542475]
textbfSDER (textbfSelf-improving textbfUnified LMMs with textbfDual stextbfElf-textbfRewards) を提案する。
論文 参考訳(メタデータ) (2025-06-09T17:38:45Z) - MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO [87.52631406241456]
近年のテキスト・ツー・イメージシステムは、マルチモーダル入力や複雑な推論タスクの処理において制限に直面している。
我々は、強化学習による推論生成を取り入れ、これらの課題に対処する統合マルチモーダルな大規模言語モデルであるMind Omniを紹介する。
論文 参考訳(メタデータ) (2025-05-19T12:17:04Z) - Harmonizing Visual Text Comprehension and Generation [31.605599298507293]
視覚テキストの理解と生成に長けた,統一的で汎用的なマルチモーダル生成モデルであるTextHarmonyを提案する。
我々は,多モード生成空間を部分的に分離して,モダリティ特化およびモダリティ非依存のLoRAエキスパートを集約するSlide-LoRAを提案する。
様々なベンチマークによる総合的な実験により,提案手法の有効性が示された。
論文 参考訳(メタデータ) (2024-07-23T10:11:56Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。