Fugu-MT 論文翻訳(概要): An Empirical Study of Mamba-based Language Models

論文の概要: An Empirical Study of Mamba-based Language Models

arxiv url: http://arxiv.org/abs/2406.07887v1
Date: Wed, 12 Jun 2024 05:25:15 GMT
ステータス: 翻訳完了
システム内更新日: 2024-06-13 18:15:17.223840
Title: An Empirical Study of Mamba-based Language Models
Title（参考訳）: マンバに基づく言語モデルに関する実証的研究
Authors: Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, Garvit Kulshreshtha, Vartika Singh, Jared Casper, Jan Kautz, Mohammad Shoeybi, Bryan Catanzaro,
Abstract要約: Mambaのような選択的な状態空間モデル(SSM)はトランスフォーマーの欠点を克服する。同じデータセット上で訓練された8B-context Mamba, Mamba-2, Transformer モデルを直接比較する。 8BのMamba-2-Hybridは、12の標準タスクで8BのTransformerを上回っている。
参考スコア（独自算出の注目度）: 69.74383762508805
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Selective state-space models (SSMs) like Mamba overcome some of the shortcomings of Transformers, such as quadratic computational complexity with sequence length and large inference-time memory requirements from the key-value cache. Moreover, recent studies have shown that SSMs can match or exceed the language modeling capabilities of Transformers, making them an attractive alternative. In a controlled setting (e.g., same data), however, studies so far have only presented small scale experiments comparing SSMs to Transformers. To understand the strengths and weaknesses of these architectures at larger scales, we present a direct comparison between 8B-parameter Mamba, Mamba-2, and Transformer models trained on the same datasets of up to 3.5T tokens. We also compare these models to a hybrid architecture consisting of 43% Mamba-2, 7% attention, and 50% MLP layers (Mamba-2-Hybrid). Using a diverse set of tasks, we answer the question of whether Mamba models can match Transformers at larger training budgets. Our results show that while pure SSMs match or exceed Transformers on many tasks, they lag behind Transformers on tasks which require strong copying or in-context learning abilities (e.g., 5-shot MMLU, Phonebook) or long-context reasoning. In contrast, we find that the 8B Mamba-2-Hybrid exceeds the 8B Transformer on all 12 standard tasks we evaluated (+2.65 points on average) and is predicted to be up to 8x faster when generating tokens at inference time. To validate long-context capabilities, we provide additional experiments evaluating variants of the Mamba-2-Hybrid and Transformer extended to support 16K, 32K, and 128K sequences. On an additional 23 long-context tasks, the hybrid model continues to closely match or exceed the Transformer on average. To enable further study, we release the checkpoints as well as the code used to train our models as part of NVIDIA's Megatron-LM project.
Abstract（参考訳）: Mambaのような選択的な状態空間モデル(SSM)は、シーケンス長の2次計算複雑性やキー値キャッシュからの大規模な推論時間メモリ要求といったトランスフォーマーの欠点を克服する。さらに、近年の研究では、SSMがトランスフォーマーの言語モデリング能力に適合または超えることが示されており、魅力的な代替手段となっている。しかし、制御された設定(例えば、同じデータ)では、これまでSSMとトランスフォーマーを比較する小さな実験しか行っていない。大規模でこれらのアーキテクチャの長所と短所を理解するため,最大3.5Tトークンのデータセットでトレーニングされた8BパラメータMamba,Mamba-2,Transformerモデルを直接比較した。また,これらのモデルを,43%のMamba-2,7%の注目,50%のMLP層(Mamba-2-Hybrid)からなるハイブリッドアーキテクチャと比較した。多様なタスクセットを使用することで、MambaモデルがTransformerとより大きなトレーニング予算で一致できるかという疑問に答える。その結果、多くのタスクにおいて、純粋なSSMはTransformerにマッチしたり、超えたりするが、強力なコピーやテキスト内学習能力(例えば、5-shot MMLU、Phonebook)や長文推論を必要とするタスクではTransformerより遅れていることがわかった。対照的に、8B Mamba-2-Hybridは、評価した12の標準タスク(平均で2.65ポイント)の8B変換器を超え、推論時にトークンを生成する場合、最大8倍高速であると予測されている。 16K,32K,128Kシーケンスをサポートするために拡張されたMamba-2-HybridおよびTransformerの変種を評価する追加実験を行った。さらに23の長いコンテキストタスクでは、ハイブリッドモデルは平均的にTransformerと密に一致または超え続けている。さらなる研究を可能にするため、NVIDIAのMegatron-LMプロジェクトの一環として、チェックポイントとモデルをトレーニングするためのコードをリリースしています。

論文の概要: An Empirical Study of Mamba-based Language Models

関連論文リスト