論文の概要: Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
- arxiv url: http://arxiv.org/abs/2609.01607v1
- Date: Tue, 01 Sep 2026 17:59:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.947257
- Title: Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
- Title(参考訳): ネイティブ統一型マルチモーダルモデルにおける理解・生成シナジーの発見:表現からタスクからシステムへ
- Authors: Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu,
- Abstract要約: 統一マルチモーダルモデル(UMM)は、単一のモデル内で視覚的理解と生成を共同で行う。
制御された、構造的にネイティブな環境で、それらの関係を表現、タスク、システムレベルで調査する。
適切な特殊化やタスク知識の共有,エンドツーエンドの最適化など,UMMの価値は統一インターフェースを超えて,共存を相乗効果に転換できることを示す。
- 参考スコア(独自算出の注目度): 90.64481646805855
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, compete for capacity, or merely coexist. We investigate their relationship at the representation, task, and system levels in a controlled, structurally native setting without pretrained vision priors. At the representation level, we find that each objective provides useful signal to the other: generation enriches the visual features learned for understanding, while understanding strengthens vision--language alignment for generation. However, when both objectives are forced through the same computation path, one tends to dominate. A task-decoupled architecture that specializes conflicting visual computation while preserving semantic interaction avoids this asymmetric degradation. At the task level, through three case studies, we find positive bidirectional transfer when understanding and generation tasks rely on shared knowledge. At the system level, we show that an end-to-end UMM outperforms a matched planner--executor pipeline on complex tasks that explicitly require both image understanding and generation. Together, these results show that the value of UMMs extends beyond a unified interface: appropriate specialization, shared task knowledge, and end-to-end optimization can turn coexistence into synergy.
- Abstract(参考訳): 統一マルチモーダルモデル(UMM)は単一のモデル内で視覚的理解と生成を共同で行うが、機能的統一は学習シナジーを保証しない。
本研究は、事前学習された視覚前処理を伴わない、制御された、構造的にネイティブな環境における、その表現、タスク、システムレベルにおけるそれらの関係について検討する。
表現レベルでは、各目的が相互に有用な信号を提供する。生成は、学習した視覚的特徴を豊かにし、理解は、生成のための視覚言語アライメントを強化する。
しかし、両方の目的が同じ計算経路を強要されると、支配的になりがちである。
意味的な相互作用を保ちながら競合する視覚計算を専門とするタスク分離アーキテクチャは、この非対称的な劣化を避ける。
タスクレベルでは、3つのケーススタディを通じて、理解と生成のタスクが共有知識に依存している場合、正の双方向転送が見つかる。
システムレベルでは、画像理解と生成の両方を明示的に要求する複雑なタスクにおいて、エンドツーエンドのUMMが一致したプランナー・エグゼクタパイプラインより優れていることを示す。
これらの結果から,UMMの価値は,適切な特殊化やタスク知識の共有,エンドツーエンドの最適化など,統一されたインターフェースを超えて拡張されることが示唆された。
関連論文リスト
- Transferability Between Understanding and Generation in Unified Multimodal Models [35.84666508890256]
Unified Multimodal Models (UMM) は、画像の理解と生成を単一のアーキテクチャに統合する。
我々は,一方のタスクにおける能力の訓練が他方のタスクにおける能力を改善するかどうかを,明示的な監督なしに検討する。
論文 参考訳(メタデータ) (2026-07-05T17:33:59Z) - Steering Visual Generation in Unified Multimodal Models with Understanding Supervision [42.765106450407814]
統一マルチモーダルモデルは、理解と生成のギャップを埋めるために考えられている。
本稿では, 個別のタスクとしてだけでなく, 生成表現を制御するための直接監督信号として, より軽量なフレームワークである「理解指向ポストトレーニング(UNO)」を提案する。
論文 参考訳(メタデータ) (2026-05-07T07:20:04Z) - Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models [98.8608163448532]
統一マルチモーダルモデル(UMM)は、視覚的理解と生成の統合において顕著な進歩を遂げた。
本稿では,UMMを教師と学生として同時に機能させる,トークンレベルの固有テキスト画像アライメント報酬機構GvUを提案する。
提案手法により,UMMの生成が大幅に向上し,視覚的理解の微粒化が促進されることを示す。
論文 参考訳(メタデータ) (2026-03-06T08:56:14Z) - Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking [154.2388970262703]
Unified Vision-Language Models (UVLM) は、単一のフレームワーク内での理解と生成の両方をサポートすることで、マルチモーダル学習を促進することを目的としている。
本稿では,解析処理と起案処理を交互に行う新たな思考パラダイムである,インターリーブド・アナライジング・ドレイティング問題解決ループ(AD-Loop)を紹介する。
テキスト思考を視覚的思考とインターリーブすることで、AD-Loopはモデルが理解と出力の両方を反復的に洗練し、真のシナジーを育むことができる。
論文 参考訳(メタデータ) (2026-02-24T23:26:09Z) - Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation [53.18286807225952]
統一マルチモーダルモデル(UMM)は、視覚的理解と生成の両方を単一のフレームワークに統合する。
単純なアーキテクチャに依存しないポストトレーニング手法であるUniMRG(Unified Multi-Representation Generation)を提案する。
提案手法は, 微粒化知覚を高め, 幻覚を低減し, 空間的理解を向上し, 同時に生成能力を向上する。
論文 参考訳(メタデータ) (2026-01-29T08:42:25Z) - Revisiting Multi-Task Visual Representation Learning [52.93947931352643]
本稿では,マルチタスク・ビジュアル事前学習フレームワークであるMTVを紹介する。
我々は、高容量の「エキスパート」モデルを利用して、高密度で構造化された擬似ラベルを大規模に合成する。
以上の結果から,MTV が "Best-of-both-worlds" のパフォーマンスを達成できることが示唆された。
論文 参考訳(メタデータ) (2026-01-20T11:59:19Z) - Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark [69.8473923357969]
統一マルチモーダルモデルは、視覚的理解と生成を共同で行うことを目的としているが、現在のベンチマークでは、その真の統合を検査することはめったにない。
提案するUni-MMMUは、8つの推論中心領域にまたがる生成と理解の双方向の相乗効果を拡大する総合的なベンチマークである。
論文 参考訳(メタデータ) (2025-10-15T17:10:35Z) - UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation [39.921363034430875]
統一された画像理解と生成は、マルチモーダル人工知能において有望なパラダイムとして浮上している。
本研究では,タスク固有の専門家モデルの理解と生成のためのモダリティアライメント行動について検討する。
タスクの干渉を避けるため,タスク固有の分岐を深いレイヤに導入しながら,タスクのタスク表現学習のための浅いレイヤを共有する,新しいY字型アーキテクチャであるUniForkを紹介した。
論文 参考訳(メタデータ) (2025-06-20T17:52:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。