論文の概要: Integrating Language Models into Listened and Imagined Speech Decoding from MEG
- arxiv url: http://arxiv.org/abs/2609.31997v1
- Date: Fri, 25 Sep 2026 20:55:26 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 03:56:11.239173
- Title: Integrating Language Models into Listened and Imagined Speech Decoding from MEG
- Title(参考訳): MEGによる音声デコーディングにおける言語モデルの統合
- Abstract要約: 我々は、MEG表現を音響および文脈言語表現と整合させるニューラルデコーダを訓練する。
音声音声の復号化は,音声音声の復号化よりも,言語モデルから恩恵を受けることが判明した。
- 参考スコア(独自算出の注目度): 19.774038251336375
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Decoding imagined speech is an important goal for brain-computer interfaces but remains challenging due to weak neural responses, low signal-to-noise ratio, and limited imagined-speech datasets. Language models provide strong contextual cues for text prediction, but how much they can help neural decoding and whether their contribution differs for decoding perceived and imagined speech remains unclear. To investigate this, we use a paired listened-imagined MEG dataset and incorporate language-model information at two stages. First, we train a contrastive neural decoder that aligns MEG representations with acoustic and contextual language representations, improving cross-subject word decoding for both listened and imagined speech. Second, at inference, we introduce a neural-constrained beam-search framework that combines neural evidence with language-model next-word probabilities. We find that imagined-speech decoding benefits more from the language model than listened-speech decoding. For Imagined speech, the best-performing balance between neural and language-model evidence shifts toward the language model, and the gain over neural-only decoding is larger. Together, these results suggest that language priors are most useful when neural evidence is weaker, making them particularly valuable for imagined-speech BCIs.
- Abstract(参考訳): 想像された音声の復号化は、脳とコンピュータのインターフェイスにとって重要な目標であるが、弱いニューラル応答、低信号対雑音比、限られた想像された音声データセットのため、依然として困難なままである。
言語モデルは、テキスト予測のための強い文脈的手がかりを提供するが、それがニューラルデコーディングにどの程度役立つか、その寄与が認識および想像された音声の復号に異なるか否かは、まだ不明である。
これを調べるために、ペア化された聴取型MEGデータセットを使用し、言語モデル情報を2段階に組み込む。
まず,MEG表現を音響および文脈言語表現に整合させるコントラスト型ニューラルデコーダを訓練し,聴取音声と想像音声の両方に対するクロスオブジェクト語デコーダを改善する。
第二に、推論時に、ニューラルエビデンスと言語モデル次の単語確率を組み合わせた、ニューラル制約されたビームサーチフレームワークを導入する。
音声音声の復号化は,音声音声の復号化よりも言語モデルから恩恵を受けることが判明した。
想像的音声では、ニューラルモデルと言語モデルのエビデンスの間で最高のパフォーマンスのバランスが言語モデルにシフトし、ニューラルのみの復号化よりも向上する。
これらの結果は、ニューラルエビデンスが弱く、特に想像された音声BCIにとって、言語先行が最も有用であることを示唆している。
関連論文リスト
- Decoding inner speech with an end-to-end brain-to-text neural interface [33.17572163528015]
音声脳-コンピュータインタフェース(BCI)は、神経活動をテキストに翻訳することで麻痺のある人々のコミュニケーションを回復することを目的としている。
本稿では、単一微分可能なニューラルネットワークを用いて、ニューラルネットワークをコヒーレントな文に変換する、エンドツーエンドのBrain-to-Textフレームワークを紹介する。
論文 参考訳(メタデータ) (2025-11-21T21:25:54Z) - Language Reconstruction with Brain Predictive Coding from fMRI Data [28.217967547268216]
予測符号化の理論は、人間の脳が将来的な単語表現を継続的に予測していることを示唆している。
textscPredFTは、BLEU-1スコアが最大27.8%$の最先端のデコード性能を実現する。
論文 参考訳(メタデータ) (2024-05-19T16:06:02Z) - SpeechAlign: Aligning Speech Generation to Human Preferences [51.684183257809075]
本稿では,言語モデルと人間の嗜好を一致させる反復的自己改善戦略であるSpeechAlignを紹介する。
我々は、SpeechAlignが分散ギャップを埋め、言語モデルの継続的自己改善を促進することができることを示す。
論文 参考訳(メタデータ) (2024-04-08T15:21:17Z) - BrainLLM: Generative Language Decoding from Brain Recordings [77.66707255697706]
本稿では,大言語モデルと意味脳デコーダの容量を利用した生成言語BCIを提案する。
提案モデルでは,視覚的・聴覚的言語刺激のセマンティック内容に整合したコヒーレントな言語系列を生成することができる。
本研究は,直接言語生成におけるBCIの活用の可能性と可能性を示すものである。
論文 参考訳(メタデータ) (2023-11-16T13:37:21Z) - Speech language models lack important brain-relevant semantics [6.626540321463248]
近年の研究では、テキストベースの言語モデルは、テキスト誘発脳活動と音声誘発脳活動の両方を驚くほど予測している。
このことは、脳内でどのような情報言語モデルが本当に予測されるのかという疑問を引き起こします。
論文 参考訳(メタデータ) (2023-11-08T13:11:48Z) - Do self-supervised speech and language models extract similar
representations as human brain? [2.390915090736061]
自己教師付き学習(SSL)によって訓練された音声と言語モデルは、音声と言語知覚の間の脳活動と強い整合性を示す。
我々は2つの代表的なSSLモデルであるWav2Vec2.0とGPT-2の脳波予測性能を評価した。
論文 参考訳(メタデータ) (2023-10-07T01:39:56Z) - Toward a realistic model of speech processing in the brain with
self-supervised learning [67.7130239674153]
生波形で訓練された自己教師型アルゴリズムは有望な候補である。
We show that Wav2Vec 2.0 learns brain-like representations with little as 600 hours of unlabelled speech。
論文 参考訳(メタデータ) (2022-06-03T17:01:46Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。