論文の概要: MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
- arxiv url: http://arxiv.org/abs/2607.21061v1
- Date: Thu, 23 Jul 2026 08:52:11 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-24 18:26:25.337499
- Title: MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
- Title(参考訳): MVEI & EmObserver:感情ステートメント判断によるMLLM指向の視覚的感情知能の強化
- Authors: Daiqing Wu, Dongbao Yang, Jiashu Yao, Hongrui Zhang, Can Ma, Yu Zhou, Sicheng Zhao,
- Abstract要約: Emotion Statement Decmentment (ESJ) は、出力を識別出力に制限しながら入力空間の表現性を保存するステートメント検証の定式化である。
ESJに最適化された感情指向MLLMであるEmObserverを、複雑なマルチステージレシピで構築する。
これらの結果は、ESJを実践的な定式化として、MVEIを総合的なベンチマークとして、EmObserverをMLLM指向の視覚的感情知性向上のための先進的なベースラインとして確立する。
- 参考スコア(独自算出の注目度): 29.753834613342274
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.
- Abstract(参考訳): AICA(Affective Image Content Analysis)は、視覚的コンテンツによって引き起こされる感情を認識し、理解することを目的としている。
しかし,Multimodal Large Language Models (MLLM) の急速な進歩にもかかわらず,その視覚的感情的知能の体系的評価は,近年のモデルリリースではほとんど行われていない。
このギャップは、従来のAICAパラダイムとMLLMのオープンエンド・インストラクション駆動性との間の構造的ミスマッチに起因し、さらなる分析により、プラウシブル応答の排除、感情分類学の制限、文脈要因の無視、労働集約アノテーションの4つの大きな制限が明らかになった。
これらの障壁を克服するために、出力を識別的判断に制約しながら入力空間の表現性を保った文の検証式である感情ステートメント判断(ESJ)を導入する。
我々はさらに、感情極性、感情解釈、シーンコンテキスト、知覚主観性にまたがる厳格に洗練されたベンチマークである、INSETS-462kを構築し、MVEIをサポートすることで、ESJを大規模にインスタンス化する労働効率の高いパイプラインであるINSETSを開発する。
評価以外にも,ESJに最適化された感情指向MLLMであるEmObserverを,精巧なマルチステージレシピで構築する。
MVEI上での広視野MLLMの広範囲な評価は、現在の人工的な視覚的感情知能に対する微妙な洞察を明らかにし、複数のAICAベンチマークでは、EmObserverの正確性、一般化、そして推論忠実さを実証している。
これらの結果は,実践的な定式化としてESJ,総合的なベンチマークとしてMVEI,MLLM指向の視覚感情知能向上のための先進的なベースラインとしてEmObserverを確立した。
コードは、https://github.com/wdqqdw/EmObserver.comでリリースされる。
関連論文リスト
- EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs [22.306059796734772]
マルチモーダルビデオにおける感情動態理解のベンチマークであるEmoTransを提案する。
EmoTransには、注意深く収集され、手動で注釈付けされたビデオクリップが1000本含まれており、12の現実世界のシナリオをカバーしている。
我々はEmoTrans上で18種類の最先端MLLMの総合的な評価を行い,2つの主な知見を得た。
論文 参考訳(メタデータ) (2026-04-25T15:10:25Z) - E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis [54.763420895859035]
脳波からの感情分析のための最初のMLLMフレームワークであるELLM2-EEG-to-Emotion Large Language Modelを提案する。
ELLMは学習可能なプロジェクション層を通じて、トレーニング済みのEEGエンコーダとQベースのLLMを統合し、マルチステージのトレーニングパイプラインを使用する。
7つの感情カテゴリーにまたがるデータセット実験により, ELLM2-EEG-to-Emotion Large Language Modelは感情分類において優れた性能を発揮することが示された。
論文 参考訳(メタデータ) (2026-01-11T13:21:20Z) - Emotion-Enhanced Multi-Task Learning with LLMs for Aspect Category Sentiment Analysis [3.605122187208041]
本稿では,感情極性とカテゴリー固有の感情を学習する,感情強化型ACSAフレームワークを提案する。
我々のアプローチは、各アスペクトカテゴリに対して、モデルが感情的な記述を生成することを可能にする。
また,Valence-Arousal-Dominance(VAD)次元フレームワークに基づく感情改善機構も導入する。
論文 参考訳(メタデータ) (2025-11-24T13:52:42Z) - Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier [53.55996102181836]
本稿では,感情関係検証器 (ERV) と説明リワードを提案する。
本手法は,対象感情と明確に一致した推論をモデルに導出する。
我々のアプローチは、説明と予測の整合性を高めるだけでなく、MLLMが感情的に一貫性があり、信頼できる対話を実現するのにも役立ちます。
論文 参考訳(メタデータ) (2025-10-27T16:40:17Z) - Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach [29.502292089901825]
この矛盾は, 既存の評価手法の制約に起因していると論じる。
これらの制約を克服する感情文判断タスクを提案する。
人間の努力を最小限に抑えて感情中心の文を効率的に構築する自動パイプラインを考案する。
論文 参考訳(メタデータ) (2025-09-26T06:30:39Z) - MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models [108.61337743051483]
MME-Emotionは,MLLMの感情的理解と推論能力の両方を評価するシステムベンチマークである。
MME-Emotionには6000以上のキュレートされたビデオクリップとタスク固有の質問回答(QA)ペアが含まれており、8つの感情的なタスクを定式化するための広いシナリオにまたがっている。
マルチエージェントシステムフレームワークを通じて分析された、感情認識と推論のためのハイブリッドメトリクスを備えた総合評価スイートが組み込まれている。
論文 参考訳(メタデータ) (2025-08-11T03:14:55Z) - EmoBench: Evaluating the Emotional Intelligence of Large Language Models [73.60839120040887]
EmoBenchは、確立された心理学理論に基づいて、マシン感情知能(EI)の包括的な定義を提案するベンチマークである。
EmoBenchには、英語と中国語で400の手作りの質問が含まれている。
以上の結果から,既存の大規模言語モデルのEIと平均的な人間の間には,かなりのギャップがみられ,今後の研究に向けての有望な方向性が浮かび上がっている。
論文 参考訳(メタデータ) (2024-02-19T11:48:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。