論文の概要: Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
- arxiv url: http://arxiv.org/abs/2609.00727v1
- Date: Tue, 01 Sep 2026 05:03:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.342547
- Title: Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
- Title(参考訳): Heard but not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
- Authors: Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj,
- Abstract要約: 本稿では,Whisper-large-v2,Qwen2-Audio-7B Instruct,Qwen2.5-Omni-7B,Chroma-4Bの4つのオープンソースモデルにおけるパラ言語情報の力学解析を行った。
我々は、中心となるカーネルアライメント、線形探索と1人の話者のアウト評価、オープンエンドトーン予測、およびコンテンツ韻律リークメトリクスを組み合わせることで、スタイル情報がオーディオエンコーダから最終出力へどのように移動するかを追跡する。
- 参考スコア(独自算出の注目度): 58.79935033105039
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, and Chroma-4B, using the Expresso dataset with controlled speaking styles. We combine centered kernel alignment, linear probing with leave one speaker out evaluation, open ended tone prediction, and a content prosody leakage metric to trace how style information moves from the audio encoder to the final output. All models strongly encode speaking style in the late encoder, that is, the top third of the audio encoder's layers, but this information is consistently degraded before reaching the output. The projector reshapes representation geometry without removing information, while decoders differ in how much style they preserve depending on architecture and training objective. At the output level, models fall into two behaviors. Some are content driven, where predictions depend mainly on text. Others are acoustic driven, where predictions vary with speaking style. The leakage metric quantifies this difference, and qualitative results confirm it. Overall, we identify a gap between what models encode and what they use, highlighting a key limitation in current audio language models.
- Abstract(参考訳): 音声言語モデルは、音声を理解するために設計されているが、その言葉以上の言葉が聞こえているかどうかは不明だ。
本稿では,4つのオープンソースモデル(Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, Chroma-4B)におけるパラ言語情報の機械的解析を行う。
我々は、中心となるカーネルアライメント、線形探索と1人の話者のアウト評価、オープンエンドトーン予測、およびコンテンツ韻律リークメトリクスを組み合わせることで、スタイル情報がオーディオエンコーダから最終出力へどのように移動するかを追跡する。
すべてのモデルは、後期エンコーダの話し方、すなわちオーディオエンコーダの3分の1の層を強くエンコードするが、この情報は出力に到達する前に一貫して劣化する。
プロジェクタは情報を取り除かずに表現幾何学を再認識するが、デコーダはアーキテクチャや訓練目的に応じてどれだけのスタイルを保存するかが異なる。
出力レベルでは、モデルは2つの行動に分類される。
一部はコンテンツ駆動型であり、予測は主にテキストに依存している。
他のものは音響駆動であり、予測は話し方によって異なる。
漏洩指標はこの差を定量化し、定性的な結果で確認する。
全体として、現在の音声言語モデルにおいて、どのモデルがエンコードされているかと何が使われているかのギャップを特定し、重要な制限を強調します。
関連論文リスト
- Testing chatbots on the creation of encoders for audio conditioned image generation [0.0]
現状の会話エージェントがCLIPテキストエンコーダを置き換える効果的なオーディオエンコーダを設計できるかどうかを検討する。
私たちは公開デモを作り、誰もがこのオーディオエンコーダを勉強して試せるようにしました。
論文 参考訳(メタデータ) (2025-09-09T18:57:10Z) - From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data [55.2480439325792]
音声対応の大規模言語モデル(ALLM)は近年,音声入力の理解と処理において大きな進歩を遂げている。
これらのモデルは典型的にはテキストベースの大規模言語モデル(LLM)に適応し、音声関連タスクのさらなるトレーニングを行う。
本研究では、現在と欠落した音を区別するALLMの能力を高めるために、コントラッシブな訓練データを生成するデータ生成フレームワークを提案する。
論文 参考訳(メタデータ) (2025-05-26T16:08:41Z) - Exploring the Role of Audio in Video Captioning [59.679122191706426]
本稿では,キャプションの音響モダリティの可能性をフル活用することを目的とした音声視覚フレームワークを提案する。
本稿では,音声とビデオ間の情報交換を改善するため,新たなローカル・グローバル融合機構を提案する。
論文 参考訳(メタデータ) (2023-06-21T20:54:52Z) - Inter-connection: Effective Connection between Pre-trained Encoder and
Decoder for Speech Translation [10.103202030679844]
本稿では,音声事前学習モデルの各層から情報を集約する相互接続機構を提案する。
この機構は, 音声事前学習モデルが凍結した場合に, パラメータを2K増加させることで, en-de, en-ja, en-zhでBLEUを約2ポイント増加させた。
論文 参考訳(メタデータ) (2023-05-26T13:01:29Z) - Autodecompose: A generative self-supervised model for semantic
decomposition [1.5990720051907859]
AutoDecomposeは、データを2つの意味的に独立した性質に分解する自己教師型生成モデルである。
音声信号にAuto Decomposeを適用し、音源(人間の声)とコンテンツを符号化する。
大規模なモデルが小さなデータセットで事前トレーニングされている場合でも,Autodecomposeはオーバーフィッティングに対して堅牢であることを示す。
論文 参考訳(メタデータ) (2023-02-06T21:18:09Z) - Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired
Speech Data [145.95460945321253]
本稿では,音響単位,すなわち擬似符号を用いたエンコーダ・デコーダネットワークのための2つの事前学習タスクを提案する。
提案したSpeech2Cは,デコーダを事前学習することなく,単語誤り率(WER)を19.2%削減できる。
論文 参考訳(メタデータ) (2022-03-31T15:33:56Z) - What all do audio transformer models hear? Probing Acoustic
Representations for Language Delivery and its Structure [64.54208910952651]
オーディオトランスフォーマーモデル mockingjay と wave2vec2.0 を比較した。
音声モデルのテキスト表面、構文、および意味的特徴に対する理解を調査します。
ネイティブ、非ネイティブ、合成、読み取り、自発的な音声データセットの完全な設定でこれを行います。
論文 参考訳(メタデータ) (2021-01-02T06:29:12Z) - Vector-quantized neural networks for acoustic unit discovery in the
ZeroSpeech 2020 challenge [26.114011076658237]
音声の離散表現を学習する問題に対処する2つのニューラルモデルを提案する。
第1モデルはベクトル量子化変分オートエンコーダ(VQ-VAE)の一種である。
第2のモデルはベクトル量子化と対比予測符号化(VQ-CPC)を組み合わせる
我々は、ZeroSpeech 2020チャレンジにおいて、英語とインドネシア語のデータをモデルとして評価した。
論文 参考訳(メタデータ) (2020-05-19T13:06:17Z) - Unsupervised Audiovisual Synthesis via Exemplar Autoencoders [59.13989658692953]
我々は,任意の個人の入力音声を,潜在的に無限に多くの出力スピーカのオーディオ視覚ストリームに変換する教師なしのアプローチを提案する。
我々は、Exemplar Autoencodersを用いて、特定のターゲット音声の音声、スタイリスティックな韻律、視覚的外観を学習する。
論文 参考訳(メタデータ) (2020-01-13T18:56:45Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。