論文の概要: LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
- arxiv url: http://arxiv.org/abs/2609.40079v1
- Date: Wed, 30 Sep 2026 16:27:25 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:28.064992
- Title: LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
- Title(参考訳): LongEmo:ロングビデオにおける感情理解と推論を目指して
- Abstract要約: LongEmoBenchは、ロングビデオの感情理解と推論に特化したベンチマークである。
LongEmoは連続したビデオストリームを処理してイベントメモリグラフを構築する。
LongEmoは最先端のパフォーマンスを実現し、イベント中心のメモリアーキテクチャの有効性を実証している。
- 参考スコア(独自算出の注目度): 33.63740775085161
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.
- Abstract(参考訳): 最近のMLLM(Multimodal Large Language Models)は、感情コンピューティングにおいて有望であることを示しているが、その推論能力は、インタラクションに制限のある短いビデオクリップに限られている。
しかし、現実の感情は単なる瞬間的な反応ではなく、過去の経験や進行中の出来事によって深く形作られた動的で累積的な過程である。
このギャップを埋めるために、長いビデオの感情理解と推論に特化したベンチマークであるLongEmoBenchを紹介します。
継続的なシーンインタラクションから複雑なエピソード開発まで、プログレッシブな能力のスケーリングを評価する。
さらに,LongEmoを提案する。LongEmoは,長期的情緒的推論の課題に対処するために設計された,メモリ拡張型エージェントフレームワークである。
LongEmoは継続的ビデオストリームを処理してEvent Memory Graphを構築する。
質問に対して、エージェントはグラフからクエリ関連イベントストリームを取得し、複数のモーダルメモリとリレーショナル依存関係を反復的に統合して最終回答を導出する。
17種類の代表的な方法の広範囲な評価は、長いビデオにおいて感情の理解と推論にかなり苦労していることを示している。
対照的にLongEmoは最先端のパフォーマンスを実現し、イベント中心のメモリアーキテクチャの有効性を実証している。
関連論文リスト
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning [66.24870234484668]
我々は,複数の相補的記憶から構築・取得する,新しいマルチモーダルメモリエージェント WorldMM を紹介する。
WorldMMは5つの長いビデオ質問回答ベンチマークで既存のベースラインを大幅に上回っている。
論文 参考訳(メタデータ) (2025-12-02T05:14:52Z) - LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction [72.19473883287948]
LongEmotionは、ロングコンテキスト感情知能(EI)タスク用に特別に設計されたベンチマークである。
感情分類、感情検出、感情QA、感情会話、感情概要、感情表現など、さまざまなタスクをカバーしている。
現実的な制約下での性能を高めるため、検索型強化世代(RAG)と協調感情モデリング(CoEM)を取り入れた。
論文 参考訳(メタデータ) (2025-09-09T05:32:45Z) - Infinite Video Understanding [50.78256932424239]
Infinite Video Understandingをブルースキー研究の目的とするフレーミングは、マルチメディアにとって重要な北の星となると我々は主張する。
我々は、この変革能力を達成するための主要な課題と研究の方向性を概説する。
論文 参考訳(メタデータ) (2025-07-11T23:07:04Z) - Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding [26.36195886824082]
Emotion-Qwenは、堅牢な感情理解と一般的な推論機能を維持するために同時に設計された統合マルチモーダルフレームワークである。
我々は,40万本以上のビデオクリップに詳細な文脈対応感情記述を付加した大規模バイリンガル・リソースであるビデオ感情推論データセットを開発した。
論文 参考訳(メタデータ) (2025-05-10T16:15:26Z) - InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions [104.90258030688256]
本研究は,ストリーミング映像とオーディオ入力とのリアルタイムインタラクションを実現するために,非絡み合いのストリーミング知覚,推論,メモリ機構を導入している。
このプロジェクトは人間のような認知をシミュレートし、多モーダルな大規模言語モデルが時間とともに継続的かつ適応的なサービスを提供できるようにする。
論文 参考訳(メタデータ) (2024-12-12T18:58:30Z) - Dilated Context Integrated Network with Cross-Modal Consensus for
Temporal Emotion Localization in Videos [128.70585652795637]
TELは、時間的行動の局所化と比較して3つのユニークな課題を提示している。
感情は時間的ダイナミクスが非常に多様である。
微粒な時間的アノテーションは複雑で、労働集約的です。
論文 参考訳(メタデータ) (2022-08-03T10:00:49Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。