論文の概要: Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
- arxiv url: http://arxiv.org/abs/2607.13188v1
- Date: Tue, 14 Jul 2026 18:39:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-16 16:39:12.566466
- Title: Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
- Title(参考訳): 同時画像理解と生成:自己補正型マルコフジャンププロセス
- Authors: Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu,
- Abstract要約: 我々は$textbfSelf-Correcting Coupled Markov Jump Processes (SC-CMJP)を紹介する。
SC-CMJPと組み合わせて、共同マルチモーダルジェネレーションのための新しいトレーニングフリーシングルパスサンプリングであるtextttCO_texttt2textttJump$を紹介する。
トレーニングと評価のために,我々は3つの大規模ジョイントマルチモーダル生成コーパスを作成し,リリースする。
- 参考スコア(独自算出の注目度): 70.61868608402723
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions $\textit{within}$ the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce $\textbf{Self-Correcting Coupled Markov Jump Processes (SC-CMJP)}$, a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce $\texttt{CO}_\texttt{2}\texttt{Jump}$ (Self-$\underline{\text{CO}}$rrecting $\underline{\text{CO}}$upled $\underline{\text{Jump}}$), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: $\text{JEdit-1M}$, $\text{JMaze-200K}$, $\text{JNono-200K}$, with matching in- and out-of-distribution benchmarks. $\texttt{CO}_\texttt{2}\texttt{Jump}$ achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling $\textit{compound}$ across the trajectory. Project page: https://coupled-jump.github.io
- Abstract(参考訳): 人間の認知は理解と生成を区別しない。
ホワイトボードの教師が話し、$\textit{together}$を描きます。
本稿では,この結合ループを人工システムに適用する。
Masked Diffusion Models (MDMs) は、このタスクに理想的には適しているが、既存のサンプルは、テキストとイメージをインターリーブまたは独立にデコードし、以前のステップ履歴のみを共有する並列ブランチで更新するが、他のモダリティの最新の決定は$\textit{within}$同じステップではない。
我々は,あるモダリティの遷移速度が他のモダリティの信頼性スコアの関数であるようなフレームワークである$\textbf{Self-Correcting Coupled Markov Jump Processes (SC-CMJP)}$を紹介する。
さらに、リマキングジャンプは、モディカルな証拠が彼らに反する瞬間を約束する。
SC-CMJPと組み合わせて、新しいマルチモーダルジェネレーション用トレーニングフリーシングルパスサンプリングである$\texttt{CO}_\texttt{2}\texttt{Jump}$ (Self-$\underline{\text{CO}}$rrecting $\underline{\text{CO}}$upled $\underline{\text{Jump}}$を紹介した。
トレーニングと評価のために、我々は3つの大規模共同マルチモーダル生成コーパスを作成し、リリースする。 $\text{JEdit-1M}$, $\text{JMaze-200K}$, $\text{Jno-200K}$。
$\texttt{CO}_\texttt{2}\texttt{Jump}$は、画像の理解と編集、そして視覚的推論(迷路やノングラムの解決)のための最高のジョイントパフォーマンスを達成する。
サンプリング器の性能はデノイングステップの数とともに単調にスケールし、その軌道を横断するクロスモーダルカップリング$\textit{compound}$の利点が証明される。
プロジェクトページ: https://coupled-jump.github.io
関連論文リスト
- Canonicalizing Multimodal Contrastive Representation Learning [76.15228959754727]
ここでは,CLIP,SigLIP,FLAVAなどのモデルファミリにおいて,埋め込み空間間の幾何学的関係が存在することを示す。
この発見は、後方互換性のあるモデルアップグレードを可能にし、コストのかかる再埋め込みを回避し、学習された表現のプライバシに影響を及ぼす。
論文 参考訳(メタデータ) (2026-02-19T18:09:36Z) - Uncovering Untapped Potential in Sample-Efficient World Model Agents [51.65485693709418]
Simulusは高度にモジュール化されたTBWMエージェントで、マルチモーダルトークン化フレームワーク、本質的なモチベーション、優先順位付けされたWMリプレイ、レグレッション・アズ・クラス化を統合している。
Simulusは3つの異なるベンチマークで、計画自由なWMに対して最先端のサンプル効率を達成する。
論文 参考訳(メタデータ) (2025-02-17T08:06:10Z) - Reasoning to Attend: Try to Understand How <SEG> Token Works [44.33848900059659]
我々は、$texttSEG>$トークンが、画像とテキストのペア内のセマンティックな類似性に寄与していることを示す。
本稿では,高活性点の誘導の下で,LMMの高強度な$textbfREA$soning機能を実現するREADを提案する。
論文 参考訳(メタデータ) (2024-12-23T17:44:05Z) - Federated Combinatorial Multi-Agent Multi-Armed Bandits [79.1700188160944]
本稿では,Banditを用いたオンライン最適化に適したフェデレーション学習フレームワークを提案する。
この設定では、エージェントのアームサブセットは、個々のアーム情報にアクセスせずにこれらのサブセットに対するノイズの多い報酬を観察し、特定の間隔で協力して情報を共有することができる。
論文 参考訳(メタデータ) (2024-05-09T17:40:09Z) - M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation [24.070247575655998]
textbf$M2Chat$は、インターリーブされたテキストイメージの会話を生成するための新しい統合マルチモーダルLLMフレームワークである。
M3Adapter$は、マルチモーダルプロンプトから、粒度の低い視覚情報と高レベルのセマンティック機能を統合する。
M3FT$ fine-tuning strategy イメージテキストアライメントとビジュアルインストラクションのために、パラメータの分離したグループを最適化する。
論文 参考訳(メタデータ) (2023-11-29T11:30:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。