論文の概要: Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
- arxiv url: http://arxiv.org/abs/2605.27772v1
- Date: Tue, 26 May 2026 23:44:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-28 17:38:55.608513
- Title: Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
- Title(参考訳): 音声 LLM は聴くか読むか? VoxParadox を用いたパラ言語障害の分析と緩和
- Authors: Jiacheng Pang, Ashutosh Chaubey, Mohammad Soleymani,
- Abstract要約: VoxParadoxは2000の検証済みのサンプルを持ち、10のパラ言語的タスクにまたがる逆のベンチマークである。
音場真理の精度は一貫して低く, 言語による回答に追従する傾向が強い。
入力プロンプトに基づいて複数のオーディオ層からの情報を適応的に結合するPrompt-Conditioned Layer Mixer (PCLM)を提案する。
- 参考スコア(独自算出の注目度): 4.088161686930475
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce VoxParadox, an adversarial benchmark with 2,000 verified examples, spanning 10 paralinguistic tasks, created with controlled speech synthesis to intentionally mismatch transcript claims and speaking style, enabling direct measurement of speech paralinguistic understanding. Evaluation of a diverse set of Audio LLMs reveals consistently low accuracy on acoustic ground truth and a strong tendency to follow language-implied (incorrect) answers. To understand the cause of this gap, we perform layer-wise probing and find that (i) paralinguistic cues can degrade in deeper encoder layers and at the encoder--LLM interface, and (ii) even when such cues are available in audio tokens, the language model frequently ignores them. To address these problems, we propose Prompt-Conditioned Layer Mixer (PCLM), which adaptively combines information from multiple audio layers based on the input prompt, and pair it with Direct Preference Optimization (DPO) to explicitly prefer acoustically supported options over language-implied alternatives. These methods substantially improve Audio LLM paralinguistic understanding, improving Audio Flamingo 3 from 17.40% to 65.20% on VoxParadox, and from 37.74% to 54.78% on MMSU paralinguistic subset. Our project page is available at https://voxparadox.github.io/.
- Abstract(参考訳): 音声大言語モデル(Audio LLMs)は、音声理解タスクにおいて強い性能を示すが、パラ言語的情報を理解する能力は限られている。
この問題を体系的に定量化するために,VoxParadoxという,2000の検証済み事例を対象とし,10のパラ言語的タスクにまたがり,制御された音声合成を用いて生成され,音声パラ言語的理解の直接的測定を可能にする。
多様なオーディオLLMの集合の評価は、音場真理に対して一貫して低い精度を示し、言語で実装された(正しくない)回答に従う傾向が強い。
このギャップの原因を理解するために、レイヤーワイドな探索を行い、それを見つける。
(i)パラ言語的キューは、より深いエンコーダ層やエンコーダ--LLMインターフェースで分解できる。
(ii) 音声トークンでそのような手がかりが利用できる場合でも、言語モデルはそれを無視することが多い。
これらの問題に対処するために、入力プロンプトに基づいて複数のオーディオ層からの情報を適応的に結合するPrompt-Conditioned Layer Mixer (PCLM)を提案する。
これらの手法はAudio LLMパラ言語理解を大幅に改善し、Audio Flamingo 3はVoxParadoxで17.40%から65.20%、MMSUパラ言語サブセットで37.74%から54.78%に改善した。
私たちのプロジェクトページはhttps://voxparadox.github.io/で公開されています。
関連論文リスト
- How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation [97.0235251827591]
大規模言語モデル (LLM) は,Large Audio Language Models (LALM) の知識バックボーンとして広く利用されている。
テキストのみの事前学習によって符号化される聴覚知識の量と、それが下流のパフォーマンスに与える影響について検討する。
その結果,家族間で聴覚知識が大きく異なり,テキストのみの結果が音響性能と強く相関していることが判明した。
論文 参考訳(メタデータ) (2026-03-19T17:50:07Z) - From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data [55.2480439325792]
音声対応の大規模言語モデル(ALLM)は近年,音声入力の理解と処理において大きな進歩を遂げている。
これらのモデルは典型的にはテキストベースの大規模言語モデル(LLM)に適応し、音声関連タスクのさらなるトレーニングを行う。
本研究では、現在と欠落した音を区別するALLMの能力を高めるために、コントラッシブな訓練データを生成するデータ生成フレームワークを提案する。
論文 参考訳(メタデータ) (2025-05-26T16:08:41Z) - Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples [55.2480439325792]
近年の音声対応大型言語モデル(ALLM)により、音声入力の処理と理解が可能になった。
これらのモデルは、しばしば既存の音響イベントを幻覚させ、現実の応用における信頼性を低下させる。
LISTENは、現在と欠落した音を識別するallMsの能力を向上するコントラスト的な訓練法である。
論文 参考訳(メタデータ) (2025-05-20T15:44:01Z) - Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models [58.43486430996411]
LALM(Large Audio-Language Models)は、最近、人間との直接の音声交換を可能にする音声対話機能をアンロックした。
オープンエンド音声対話理解におけるLALMの性能を評価するための音声対話理解ベンチマーク(ADU-Bench)を提案する。
ADU-Benchには、LALMの評価のための2万以上のオープンエンドオーディオダイアログが含まれている。
論文 参考訳(メタデータ) (2024-12-06T16:34:15Z) - Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models [0.9285295512807729]
AQA(Audio Question Answering)タスクには、オーディオイベント分類、オーディオキャプション、オープンエンド推論が含まれる。
LALMは一般的な音声理解では優れているが、時間的推論では限られている。
本稿では,音声時間的推論におけるこれらの課題と限界について述べる。
論文 参考訳(メタデータ) (2024-09-10T05:26:53Z) - Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities [37.02115473120654]
音声を理解するために大きな言語モデル(LLM)を拡張することは、様々な現実世界のアプリケーションにとって非常に重要である。
本稿では,1)強音声理解能力を備えた新しい音声言語モデルであるAudio Flamingoを提案する。
論文 参考訳(メタデータ) (2024-02-02T18:58:34Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。