論文の概要: From Neurons to Conversation: Speech Brain-Computer Interfaces
- arxiv url: http://arxiv.org/abs/2609.36736v1
- Date: Tue, 29 Sep 2026 05:10:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.202154
- Title: From Neurons to Conversation: Speech Brain-Computer Interfaces
- Title(参考訳): ニューロンから会話へ:音声脳-コンピュータインタフェース
- Abstract要約: 音声脳-コンピュータインタフェース(BCI)は、音声、言語、またはコミュニケーション意図に関連する神経活動から、テキスト、合成音声、アバター制御などの外部出力に変換することによって、コミュニケーションを回復することを目的としている。
近年, 皮質内および皮質内記録, ディープシーケンスモデル, 言語モデル支援復号法が進歩し, 急速に進歩している。
これらの成果は、音声BCIが単にニューラル・トゥ・テキスト・デコーダではないことも明らかにしている。それらは、ニューラル表現、記録ハードウェア、デコードアーキテクチャ、言語優先、フィードバック、ユーザー学習が時間とともに対話する適応型臨床システムである。
- 参考スコア(独自算出の注目度): 0.9381376621526817
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.
- Abstract(参考訳): 音声脳-コンピュータインタフェース(BCI)は、音声、言語、またはコミュニケーション意図に関連する神経活動から、テキスト、合成音声、アバター制御などの外部出力に変換することによって、コミュニケーションを回復することを目的としている。
近年の皮質内および皮質内記録、ディープシーケンスモデル、言語モデルによるデコーディングは、高速な音声復号化や、より自然主義的な音声合成など、急速に進歩している。
しかし、これらの成果は、音声BCIが単にニューラル・トゥ・テキスト・デコーダではないことも明らかにしている。
それらは適応的な臨床システムであり、神経表現、記録ハードウェア、デコードアーキテクチャ、言語優先、フィードバック、ユーザー学習が時間とともに相互作用する。
ここでは,システムレベルの観点から,音声BCI研究を合成する。
まず、音声と言語の神経基質を調べ、その階層的、分散的、時間的構造的、非定常的な組織を強調した。
次に,録音および復号化の選択,閉ループ適応,評価,臨床翻訳,倫理について検討する。
これらの領域全体では、信号分解能と侵襲性、低レベルモーターと高レベルセマンティックターゲット、デコーダの精度とユーザエージェンシー、言語モデル流布と忠実なニューラルエビデンスとのトレードオフが繰り返されている。
次世代の音声BCIは、オフラインの精度だけでなく、セッション間の堅牢性、キャリブレーションの負担、レイテンシ、不確実性、ユーザビリティ、意図しない復号化に対する保護によって評価されるべきである。
音声BCIを適応的でユーザ中心のシステムとして再定義することにより、音声神経科学、神経工学、機械学習、臨床実践、神経倫理学にまたがる学際的な優先順位を概説する。
関連論文リスト
- Integrating Language Models into Listened and Imagined Speech Decoding from MEG [19.774038251336375]
我々は、MEG表現を音響および文脈言語表現と整合させるニューラルデコーダを訓練する。
音声音声の復号化は,音声音声の復号化よりも,言語モデルから恩恵を受けることが判明した。
論文 参考訳(メタデータ) (2026-09-25T20:55:26Z) - Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication [45.424817836500175]
本研究は,様々な音声モードにおける未確認文に対する音声合成の可能性について検討する。
本研究では,高密度脳波(EEG)信号から抽出した音素レベル情報と筋電図(EMG)信号とを独立に利用した。
本研究は, 生体信号に基づく文レベルの音声合成が未確認文の再構成に有効であることを示すものである。
論文 参考訳(メタデータ) (2025-10-31T07:31:13Z) - sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment [8.466223794246261]
本稿では,凍結したCLIPモデルの文埋め込み空間に単射ステレオ脳波信号(sEEG)を投影するコントラスト学習フレームワークであるSSENSEを提案する。
本手法は,自然主義映画視聴データセットから,時系列のsEEGと音声の書き起こしについて評価する。
論文 参考訳(メタデータ) (2025-04-20T03:01:42Z) - Decoding Continuous Character-based Language from Non-invasive Brain Recordings [33.11373366800627]
本研究では,単心的非侵襲的fMRI記録から連続言語を復号する手法を提案する。
文字ベースのデコーダは、固有の文字構造を特徴とする連続言語の意味的再構成のために設計されている。
被験者間での単一の試行から連続言語を復号化できることは、非侵襲的な言語脳-コンピュータインタフェースの有望な応用を実証している。
論文 参考訳(メタデータ) (2024-03-17T12:12:33Z) - BrainLLM: Generative Language Decoding from Brain Recordings [77.66707255697706]
本稿では,大言語モデルと意味脳デコーダの容量を利用した生成言語BCIを提案する。
提案モデルでは,視覚的・聴覚的言語刺激のセマンティック内容に整合したコヒーレントな言語系列を生成することができる。
本研究は,直接言語生成におけるBCIの活用の可能性と可能性を示すものである。
論文 参考訳(メタデータ) (2023-11-16T13:37:21Z) - Toward a realistic model of speech processing in the brain with
self-supervised learning [67.7130239674153]
生波形で訓練された自己教師型アルゴリズムは有望な候補である。
We show that Wav2Vec 2.0 learns brain-like representations with little as 600 hours of unlabelled speech。
論文 参考訳(メタデータ) (2022-06-03T17:01:46Z) - Open Vocabulary Electroencephalography-To-Text Decoding and Zero-shot
Sentiment Classification [78.120927891455]
最先端のブレイン・トゥ・テキストシステムは、ニューラルネットワークを使用して脳信号から直接言語を復号することに成功した。
本稿では,自然読解課題における語彙的脳波(EEG)-テキスト列列列復号化とゼロショット文感性分類に問題を拡張する。
脳波-テキストデコーディングで40.1%のBLEU-1スコア、ゼロショット脳波に基づく3次感情分類で55.6%のF1スコアを達成し、教師付きベースラインを著しく上回る結果となった。
論文 参考訳(メタデータ) (2021-12-05T21:57:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。