論文の概要: Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
- arxiv url: http://arxiv.org/abs/2606.26493v2
- Date: Mon, 29 Jun 2026 20:21:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-01 13:50:27.698207
- Title: Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context
- Title(参考訳): Nemotron-Labs-TwoTower:事前制約付き自己回帰文脈を用いた拡散言語モデリング
- Authors: Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro,
- Abstract要約: TwoTowerはブロックワイドの自己回帰拡散モデルで、役割を2つの塔に分離する。
Nemotron-3-Nano-30B-A3Bは30BハイブリッドのMamba-Transformer MoEモデルで、約2.1Tトークンで訓練されている。
- 参考スコア(独自算出の注目度): 44.363343506907704
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a single network for both context representation and iterative denoising, forcing one model to serve both roles and limiting its capacity for either role. We propose TwoTower, a block-wise autoregressive diffusion model that decouples these roles into two towers: a frozen AR context tower that causally processes clean tokens, and a trainable diffusion denoiser tower with bidirectional block attention that refines noisy blocks via cross-attention to the context. Built on Nemotron-3-Nano-30B-A3B, an open-weight 30B hybrid Mamba-Transformer MoE model, and trained on approximately 2.1T tokens, Nemotron-Labs-TwoTower retains 98.7% of the autoregressive baseline's quality while offering 2.42X higher wall-clock generation throughput. We release the code and model weights at https://huggingface.co/collections/nvidia/nemotron-labs-twotower.
- Abstract(参考訳): 拡散言語モデルは、並列および反復生成の可能性のため、自己回帰モデルに代わる有望な代替手段を提供する。
しかし、既存のアプローチでは、コンテキスト表現と反復的記述の両方に単一のネットワークを使用し、1つのモデルが両方の役割を果たせ、それぞれの役割の能力を制限する。
本研究では,これらの役割を2つのタワーに分離するブロックワイド自己回帰拡散モデルであるTwoTowerを提案する。
Nemotron-3-Nano-30B-A3Bは30BハイブリッドのMamba-Transformer MoEモデルであり、約2.1Tトークンで訓練されている。
コードとモデルの重み付けはhttps://huggingface.co/collections/nvidia/nemotron-labs-twotowerで公開しています。
関連論文リスト
- Breaking the Bottleneck with DiffuApriel: High-Throughput Diffusion LMs with Mamba Backbone [6.76700377196741]
両方向マンバのバックボーン上に構築されたマスク付き拡散言語モデルであるDiffuAprielを紹介する。
この結果から, 双方向状態空間アーキテクチャは, マスク拡散LMの強力なデノイザとして機能することが示唆された。
論文 参考訳(メタデータ) (2025-11-19T23:23:49Z) - The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation [53.837937703425794]
LanDiffは、自己回帰言語モデルと拡散モデルの強みを相乗化するハイブリッドフレームワークである。
本アーキテクチャでは,(1)効率的なセマンティック圧縮により3次元視覚特徴をコンパクトな1次元表現に圧縮するセマンティック・トークンー,(2)高レベルのセマンティックな関係を持つセマンティック・トークンを生成する言語モデル,(3)粗いセマンティクスを高忠実なビデオに洗練するストリーミング拡散モデルを紹介する。
論文 参考訳(メタデータ) (2025-03-06T16:53:14Z) - Energy-Based Diffusion Language Models for Text Generation [126.23425882687195]
エネルギーベース拡散言語モデル(Energy-based Diffusion Language Model, EDLM)は、拡散ステップごとに全シーケンスレベルで動作するエネルギーベースモデルである。
我々のフレームワークは、既存の拡散モデルよりも1.3$times$のサンプリングスピードアップを提供する。
論文 参考訳(メタデータ) (2024-10-28T17:25:56Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。