論文の概要: BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
- arxiv url: http://arxiv.org/abs/2608.04156v1
- Date: Tue, 04 Aug 2026 19:07:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.599836
- Title: BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
- Title(参考訳): BrainBench: 包括的なEEG理解のための大規模言語モデルのベンチマーク
- Authors: Yangxuan Zhou, Sha Zhao, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan,
- Abstract要約: 包括的で命令条件のEEG理解のための統合ベンチマークであるベンチマークネームを導入する。
4つのサブセットで構成されており、-Foundational Analysis、Sleep Assessment、Neurocognitive Assessment、Physological Integrationである。
結果はモデル、サブセット、難易度、実行パラダイムによって大きく異なり、EEGの能力がモデルとその運用に依存していることを示している。
- 参考スコア(独自算出の注目度): 29.151563019741545
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.
- Abstract(参考訳): 脳波解析(EEG)解析は、予め定義されたラベルを録音に割り当てるだけでなく、自然言語命令、信号処理、量的証拠、科学的解釈を接続するワークフローを必要とする。
我々はこの能力を「emph{comprehensive EEG understanding}」と呼ぶ。
しかし、既存の評価は、主に独立したデコードタスクやシステム固有のデモをターゲットとしており、大きな言語モデル(LLM)の能力は十分に定量化されていない。
包括的、命令条件付きEEG理解のための統一ベンチマークである \benchmarkname{} を紹介する。
それは、-Foundational Analysis、Sleep Assessment、Neurocognitive Assessment、Physological Integrationの4つのサブセットで構成されている。-17データセット、 \numcases{}タスク、 \numinstances{}リアルタイムインスタンスをカバーしている。
任意の生理的信号を持つ脳波記録と指示が与えられた場合、システムは分析を行い、科学的に根拠づけられたレポートを生成し、必要に応じて人工物を生成する必要がある。
出力は数値、分類、集合、シーケンス、意味、アーティファクト検証を通じて評価される。
我々は、CodeActによる自律的なコード実行とBrainAgentによる構造化されたエージェント分析という、2つのパラダイムの下で、100K以上の実行にまたがる代表LSMを評価した。
結果はモデル、サブセット、難易度、実行パラダイムによって大きく異なり、EEGの能力がモデルとその運用に依存していることを示している。
\benchmarkname{} は LLM ベースの EEG 理解を促進する再現可能なテストベッドを提供する。
コードとベンチマークはまもなくリリースされ、評価結果は継続的に更新される。
関連論文リスト
- CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification [3.991518853474585]
MNE-Pythonを基盤とした認知脳波分析エージェントであるCogEEGAgentを提案する。
EEG固有の科学的ハーネスは、意味を科学的権威から分離する。
科学的エージェントは、フレキシブル言語理解とフェールクロースされた推論とリリースの制御を組み合わせられるかを示す。
論文 参考訳(メタデータ) (2026-07-27T20:09:05Z) - OmniEEG-Bench: A Standardized Evaluation Benchmark for EEG Foundation Models [22.964421663748755]
我々は,脳波基礎モデル(FM)のための統一ベンチマークとダウンストリームタスクロードマップであるOmniEEG-Benchを紹介する。
脳波FMの評価を、(i)信号信頼性、(ii)生体計測と疾患、(iii)意識と状態、(iv)認知と感情、(v)自然主義的刺激復号、(vi)運動と相互作用の6つのタスクファミリーに分類する。
代表的なEEGファンデーションモデル10をベンチマークし、さまざまな評価設定をカバーするリーダーボードを報告します。
論文 参考訳(メタデータ) (2026-05-30T17:20:04Z) - E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis [54.763420895859035]
脳波からの感情分析のための最初のMLLMフレームワークであるELLM2-EEG-to-Emotion Large Language Modelを提案する。
ELLMは学習可能なプロジェクション層を通じて、トレーニング済みのEEGエンコーダとQベースのLLMを統合し、マルチステージのトレーニングパイプラインを使用する。
7つの感情カテゴリーにまたがるデータセット実験により, ELLM2-EEG-to-Emotion Large Language Modelは感情分類において優れた性能を発揮することが示された。
論文 参考訳(メタデータ) (2026-01-11T13:21:20Z) - AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment [69.06977852423564]
画像品質評価(IQA)は、人間の視覚系に根ざした知覚品質の定量化と解釈の両方を反映している。
AgenticIQAは、IQAを歪み検出、歪み解析、ツール選択、ツール実行の4つのサブタスクに分解する。
本稿では,IQAエージェントに適した大規模命令データセットであるAgenticIQA-200Kと,VLMベースのIQAエージェントの計画,実行,要約機能を評価するための最初のベンチマークであるAgenticIQA-Evalを紹介する。
論文 参考訳(メタデータ) (2025-09-30T09:37:01Z) - ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying [29.301155351226765]
本稿では,セマンティックタスク命令とクエリによるEEG-Language Alignmentの基礎モデルであるELASTIQを提案する。
ELASTIQは、タスク認識のセマンティックガイダンスを統合し、構造化され言語的に整合したEEG埋め込みを生成する。
運動画像,感情認識,定常的視覚誘発電位,隠蔽音声,医療タスクを対象とする20のデータセット上でELASTIQを評価した。
論文 参考訳(メタデータ) (2025-09-29T05:29:12Z) - WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities [55.00677513249723]
脳波信号は認知過程と固有の神経状態の両方を同時に符号化する。
我々は、EEG信号とその対応するモダリティを統一意味空間にマッピングし、一般化された解釈を実現する。
結果として得られたモデルは、柔軟でオープンな会話をサポートしながら、堅牢な分類精度を示す。
論文 参考訳(メタデータ) (2025-09-26T06:21:51Z) - IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis [60.32962597618861]
IDA-Benchは、多ラウンドの対話シナリオで大規模言語モデルを評価する新しいベンチマークである。
エージェント性能は、最終的な数値出力と人間由来のベースラインを比較して判断する。
最先端のコーディングエージェント(Claude-3.7-thinkingなど)でさえ50%のタスクを成功させ、シングルターンテストでは明らかでない制限を強調している。
論文 参考訳(メタデータ) (2025-05-23T09:37:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。