論文の概要: Unified Audio Intelligence Without Regressing on Text Intelligence
- arxiv url: http://arxiv.org/abs/2607.05196v1
- Date: Mon, 06 Jul 2026 15:11:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:30.205102
- Title: Unified Audio Intelligence Without Regressing on Text Intelligence
- Title(参考訳): テキストインテリジェンスに依存しない統一音声インテリジェンス
- Abstract要約: 我々は,統一音声テキストLLMであるNemotron-Labs-Audex-30B-A3B(Audex)を紹介する。
Audexは単一のTransformerデコーダで単純な統一設計を採用する。
Audexは最先端の音声理解、音声認識と翻訳、テキスト音声合成、音声生成、音声音声生成を提供する。
- 参考スコア(独自算出の注目度): 86.55418188552768
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.
- Abstract(参考訳): オーディオインテリジェンスは、音声と音声の両方を理解し、推論し、生成する。
本研究では,Nemotron-Labs-Audex-30B-A3B(Audex)を紹介する。
オーディオ入力はエンコードされ、テキスト埋め込み空間に投影され、テキストトークンと量子化されたオーディオ出力トークンは生成時に一様に扱われる。
このアーキテクチャは、強力な音声テキスト融合、シームレスなマルチモーダル生成、標準LLMトレーニングと推論インフラとの互換性を実現する。
トレーニングでは、157.4Bの音声トークンと320.5Bのテキストトークンからなるオーディオテキストデータセットを慎重にキュレートする。
我々はこれらのデータセットに多段階教師あり訓練を適用し、次いでテキストのみのカスケードRLとマルチドメインのオンライン蒸留を行った。
Audexは、最先端の音声理解、音声認識と翻訳、テキスト音声合成、音声生成、音声音声生成を提供すると同時に、テキストのみのLLMバックボーンの非常に説得力のある推論、アライメント、知識、長文、エージェント機能を、限界あるいは非回帰で保持する。
オープンな研究を容易にするためのモデルチェックポイントをリリースする。
関連論文リスト
- From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data [55.2480439325792]
音声対応の大規模言語モデル(ALLM)は近年,音声入力の理解と処理において大きな進歩を遂げている。
これらのモデルは典型的にはテキストベースの大規模言語モデル(LLM)に適応し、音声関連タスクのさらなるトレーニングを行う。
本研究では、現在と欠落した音を区別するALLMの能力を高めるために、コントラッシブな訓練データを生成するデータ生成フレームワークを提案する。
論文 参考訳(メタデータ) (2025-05-26T16:08:41Z) - Probing Audio-Generation Capabilities of Text-Based Language Models [5.4211188445379825]
本研究では,大規模言語モデルが音声を生成できる範囲について検討する。
我々は、音声生成の複雑さを徐々に増大させる3層アプローチを採用する。
以上の結果から,LLMは基本的音声特徴を生成できるが,音声の複雑さが増すにつれて性能が低下することが明らかとなった。
論文 参考訳(メタデータ) (2025-05-04T23:46:01Z) - LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT [65.69648099999439]
Generative Pre-trained Transformer (GPT) モデルは、様々な自然言語処理タスクにおいて顕著なパフォーマンスを実現している。
音声認識, 理解, 生成のための新しい音声・テキストGPTベースのLLMであるLauraGPTを提案する。
論文 参考訳(メタデータ) (2023-10-07T03:17:59Z) - Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM [19.36630667212398]
本稿では,事前学習された大規模言語モデル(LLM)を適応させて,音声質問応答(QA)と音声継続を行う新しいアプローチであるSpectronを提案する。
我々のアプローチの鍵は、音声認識、テキスト継続、音声合成を共同で監督する訓練目標である。
提案手法は話者保存とセマンティック・コヒーレンスにおいて既存の言語モデルを上回る。
論文 参考訳(メタデータ) (2023-05-24T15:39:43Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。