論文の概要: JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning
- arxiv url: http://arxiv.org/abs/2605.13013v1
- Date: Wed, 13 May 2026 05:07:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-14 23:30:27.822225
- Title: JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning
- Title(参考訳): JEDI:オンラインモデルに基づく強化学習のための統合埋め込み拡散世界モデル
- Authors: Jing Yu Lim, Rushi Shah, Zarif Ikram, Samson Yu, Haozhe Ma, Tze-Yun Leong, Dianbo Liu,
- Abstract要約: Joint Embedding Diffusion (JEDI)は、世界初のオンラインエンドツーエンドの潜伏拡散モデルである。
JEDIはAtari100kで競争力があり、直接に比較して訓練された潜伏者でベースラインを上回ります。
- 参考スコア(独自算出の注目度): 6.0332038912905555
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Diffusion world models have recently become competitive for online model-based reinforcement learning, but current approaches expose a tension: pixel diffusion is effective but computationally expensive while the latest latent diffusion approach improves efficiency yet performs subpar. The latter also relies on separately trained latents rather than the end-to-end world-model objectives that have driven much of modern MBRL progress. In particular, JEPA-style predictive representation learning has emerged as an especially promising direction for world modeling and MBRL. Concurrently, diffusion-style objectives have gained traction across multiple domains, with iterative refinement as a promising approach for multimodal and stochastic targets. Taken together, these trends motivate Joint Embedding DIffusion (JEDI), the first online end-to-end latent diffusion world model. JEDI learns its latent space directly from the diffusion denoising loss with a JEPA framework, using denoising to learn and predict future latents rather than relying on reconstruction and pretrained models. We provide a theoretical motivation showing that conventional JEPA objectives induce a predictive information bottleneck, and that conditional diffusion denoising admits a closely related predictive-compression decomposition. Empirically, JEDI is competitive on Atari100k and outperforms the baseline with seperately trained latents where directly comparable. Relative to the pixel diffusion baseline, JEDI uses 43% less VRAM, over 3$\times$ faster world-model sampling, and 2.5$\times$ faster training. JEDI also exhibits a markedly different task-level performance profile from the pixel baseline, suggesting that end-to-end predictive latents change more than compute alone.
- Abstract(参考訳): 拡散世界モデルは近年、オンラインモデルに基づく強化学習において競争力を持つようになったが、現在のアプローチでは緊張が浮き彫りになっている。
後者は、現代のMBRLの進歩の多くを駆動するエンド・ツー・エンドの世界モデル目標よりも、個別に訓練された潜伏者に依存している。
特にJEPAスタイルの予測表現学習は世界モデリングとMBRLにとって特に有望な方向として現れている。
同時に、拡散スタイルの目的は複数の領域にまたがって勢いを増し、反復的洗練はマルチモーダルおよび確率的対象に対する有望なアプローチである。
これらの傾向は、最初のオンラインエンドツーエンドの潜伏拡散モデルであるJEDI(Joint Embedding Diffusion)を動機付けている。
JEDIはJEPAフレームワークによる拡散デノゲーション損失から直接潜伏空間を学習し、デノゲーションを使用して、再構築や事前訓練されたモデルに頼るのではなく、将来の潜伏者を学習し、予測する。
本稿では,従来のJEPAの目的が予測情報のボトルネックを生じさせ,条件拡散分極が密接に関連する予測圧縮分解を許容することを示す理論的動機を与える。
経験的に、JEDIはAtari100kで競争力があり、直接に匹敵する訓練を受けた潜伏者でベースラインを上回っている。
ピクセル拡散ベースラインとは対照的に、JEDIはVRAMを43%削減し、3$\times$より高速なワールドモデルサンプリング、2.5$\times$より高速なトレーニングを使用する。
JEDIはまた、ピクセルベースラインと明らかに異なるタスクレベルのパフォーマンスプロファイルを示しており、エンドツーエンドの予測ラテントが計算だけでなく変化することを示唆している。
関連論文リスト
- Deep Leakage with Generative Flow Matching Denoiser [54.05993847488204]
再建プロセスに先立って生成フローマッチング(FM)を組み込んだ新しい深部リーク攻撃(DL)を導入する。
当社のアプローチは、ピクセルレベル、知覚的、特徴に基づく類似度測定において、最先端の攻撃よりも一貫して優れています。
論文 参考訳(メタデータ) (2026-01-21T14:51:01Z) - Consistent World Models via Foresight Diffusion [56.45012929930605]
我々は、一貫した拡散に基づく世界モデルを学習する上で重要なボトルネックは、最適下予測能力にあると主張している。
本稿では,拡散に基づく世界モデリングフレームワークであるForesight Diffusion(ForeDiff)を提案する。
論文 参考訳(メタデータ) (2025-05-22T10:01:59Z) - Generalized Interpolating Discrete Diffusion [65.74168524007484]
仮面拡散はその単純さと有効性のために一般的な選択である。
ノイズ発生過程の設計において、より柔軟性の高い離散拡散(GIDD)を補間する新しいファミリを一般化する。
GIDDの柔軟性をエクスプロイトし、マスクと均一ノイズを組み合わせたハイブリッドアプローチを探索し、サンプル品質を向上する。
論文 参考訳(メタデータ) (2025-03-06T14:30:55Z) - Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion [55.185588994883226]
VQ-LCMDは、学習を安定させる埋め込み空間内の連続空間潜在拡散フレームワークである。
VQ-LCMDは、関節埋め込み拡散変動下界と整合整合性(CM)損失を組み合わせた新しいトレーニング目標を使用する。
実験により,提案したVQ-LCMDは離散状態潜伏拡散モデルと比較して,FFHQ,LSUN教会,LSUNベッドルームにおいて優れた結果が得られることが示された。
論文 参考訳(メタデータ) (2024-10-18T09:12:33Z) - Denoising with a Joint-Embedding Predictive Architecture [21.42513407755273]
私たちはD-JEPA(Joint-Embedding Predictive Architecture)でDenoisingを紹介します。
本稿では,JEPAをマスク画像モデリングの一形態として認識することにより,一般化した次世代予測戦略として再解釈する。
また,拡散損失を利用して確率分布をモデル化し,連続空間におけるデータ生成を可能にする。
論文 参考訳(メタデータ) (2024-10-02T05:57:10Z) - Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation [59.184980778643464]
ファインチューニング拡散モデル : 生成人工知能(GenAI)の最前線
本稿では,拡散モデル(SPIN-Diffusion)のための自己演奏ファインチューニングという革新的な手法を紹介する。
提案手法は従来の教師付き微調整とRL戦略の代替として,モデル性能とアライメントの両方を大幅に改善する。
論文 参考訳(メタデータ) (2024-02-15T18:59:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。