論文の概要: SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations
- arxiv url: http://arxiv.org/abs/2607.20445v1
- Date: Wed, 13 May 2026 13:00:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-27 00:46:13.164783
- Title: SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations
- Title(参考訳): SCoPE:会話における感情認識のためのシフト対応話者記述型先行者
- Abstract要約: SCoPE(Speaker-Conditioned Priors over Emotions)を紹介する。
SCoPEは、各話者の感情履歴を利用する軽量モジュールである。
我々は、SCoPEとマルチモーダルエビデンスからの事前のバランスをモデルに導くために、感情シフト予測を組み込んだ。
- 参考スコア(独自算出の注目度): 9.535514750203266
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly between contrasting emotions such as happiness and anger. Instead, emotions tend to evolve smoothly, and these patterns are often speaker-specific. Some people might escalate, while others gradually cool down over time. Furthermore, when emotions change during a conversation, they are often driven by contextual factors, such as newly received information or unexpected events. Even though progress has been made in Emotion Recognition in Conversations (ERC), most existing approaches still rely heavily on overt evidence and do not sufficiently model these non-apparent factors. Especially in multimodal settings, this makes these models fragile when the signals are noisy (e.g., occluded faces, slang expressions, or microphone noise). To address these limitations, we introduce Speaker-Conditioned Priors over Emotions (SCoPE). SCoPE is a light weight module that utilizes the emotional history of each speaker and explicitly models their priors for use in subsequent emotion classification. Second, we incorporate emotion shift prediction, a well-established concept in ERC, to guide the model in balancing the priors from SCoPE and multimodal evidence. Finally, we propose a shift-aware fusion mechanism that performs precision-weighted logit integration between multimodal evidence and the speaker prior, forming a Bayesian-inspired product-of-experts formulation. This dynamic fusion allows the model to rely on historical priors when emotions persist and to prioritize multimodal evidence when shifts are likely. Experimental results show our model achieves superior performance over recent state-of-the-art models on the IEMOCAP dataset in multimodal settings.
- Abstract(参考訳): 会話では、人間の感情は過渡的であるが、複数の発話にまたがって持続する傾向がある。
例えば、幸福や怒りといった対照的な感情を瞬時に切り替えることは滅多にありません。
その代わり、感情は滑らかに進化する傾向があり、これらのパターンはしばしば話者固有のものである。
ある人はエスカレートするかもしれないが、ある人は徐々に冷えていく。
さらに、会話中に感情が変化する際には、しばしば、新たに受信された情報や予期せぬ出来事などの文脈的要因によって引き起こされる。
会話における感情認識(Emotion Recognition in Conversations, ERC)は進歩しているが、既存のアプローチの多くは未だに過剰な証拠に大きく依存しており、これらの不明瞭な要因を十分にモデル化していない。
特にマルチモーダル環境では、これらのモデルはノイズ(例えば、隠蔽された顔、スラング表現、マイクロホンノイズ)のときに壊れやすい。
これらの制約に対処するため、我々はScoPE(Speaker-Conditioned Priors over Emotions)を導入する。
SCoPEは、各話者の感情履歴を生かした軽量モジュールであり、その後の感情分類に使用するためにそれらの先行を明示的にモデル化する。
第2に,感情変化予測をERCで確立した概念として組み込んで,SCoPEとマルチモーダルエビデンスから事前のバランスをとるモデルを導出する。
最後に,マルチモーダルエビデンスと話者との高精度なロジット統合を実現するシフトアウェア融合機構を提案する。
このダイナミックフュージョンは、感情が持続する過去の過去の履歴をモデルが頼りにし、シフトの可能性が高い時にマルチモーダルな証拠を優先順位付けすることを可能にする。
実験結果から,IEMOCAPデータセットのマルチモーダル環境での最近の最先端モデルよりも優れた性能が得られた。
関連論文リスト
- AffectVerse: Emotional World Models for Multimodal Affective Computing [56.144242722718985]
AffectVerseは、短時間の潜伏感情予測のためのアクションフリー表現レベルモジュールである。
EWMには3つのモジュールが含まれている。 1) クロスモーダルなテンポラル・イマジネーションは、複数ステップのロールアウトで過去のトークンから将来のビデオ/オーディオ表現を予測する。
EWMは想像されたトークンをモダリティ対応の信仰トークンに圧縮する。
これらの結果は、予測的信念状態モデリングが感情コンピューティングの実用的な代替手段であることを示唆している。
論文 参考訳(メタデータ) (2026-05-19T15:05:00Z) - On the Emotion Understanding of Synthesized Speech [63.13411068766772]
感情は音声対話における中核的なパラ言語的特徴である。
現在の音声感情認識(SER)モデルは、合成音声に一般化できない。
生成音声言語モデル(SLM)は、パラ言語的手がかりを無視しながら、テキスト意味論から感情を推測する傾向がある。
論文 参考訳(メタデータ) (2026-03-17T13:11:14Z) - ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning [67.22219034602514]
ADEPT(Agentic Decoding of Emotion via Evidence Probing Tools)は,感情認識をマルチターン探索プロセスとして再構成するフレームワークである。
ADEPTはSLLMを進化する候補感情を維持するエージェントに変換し、専用のセマンティックおよび音響探査ツールを適応的に呼び出す。
ADEPTは、ほとんどの設定において主感情の精度を向上し、微妙な感情の特徴を著しく改善することを示した。
論文 参考訳(メタデータ) (2026-02-13T08:33:37Z) - GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations [35.63053777817013]
GatedxLSTMは、会話におけるマルチモーダル感情認識(ERC)モデルである。
話者と会話相手の双方の声と書き起こしを考慮し、感情的なシフトを駆動する最も影響力のある文章を特定する。
4クラスの感情分類において,オープンソース手法間でのSOTA(State-of-the-art)性能を実現する。
論文 参考訳(メタデータ) (2025-03-26T18:46:18Z) - Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion [87.18073195745914]
人間の感情が感情の予測において有意であると考えられる特徴とどのように相関するかを検討する。
EmoTriggerを用いて、感情のトリガーを識別する大規模言語モデルの能力を評価する。
分析の結果、感情のトリガーは感情予測モデルにとって健全な特徴ではなく、様々な特徴と感情検出のタスクの間に複雑な相互作用があることが判明した。
論文 参考訳(メタデータ) (2023-11-16T06:20:13Z) - Dynamic Causal Disentanglement Model for Dialogue Emotion Detection [77.96255121683011]
隠れ変数分離に基づく動的因果解離モデルを提案する。
このモデルは、対話の内容を効果的に分解し、感情の時間的蓄積を調べる。
具体的には,発話と隠れ変数の伝搬を推定する動的時間的ゆがみモデルを提案する。
論文 参考訳(メタデータ) (2023-09-13T12:58:09Z) - GM-TCNet: Gated Multi-scale Temporal Convolutional Network using Emotion
Causality for Speech Emotion Recognition [14.700043991797537]
本稿では,新しい感情的因果表現学習コンポーネントを構築するために,GM-TCNet(Gated Multi-scale Temporal Convolutional Network)を提案する。
GM-TCNetは、時間領域全体の感情のダイナミクスを捉えるために、新しい感情因果表現学習コンポーネントをデプロイする。
我々のモデルは、最先端技術と比較して、ほとんどのケースで最高の性能を維持している。
論文 参考訳(メタデータ) (2022-10-28T02:00:40Z) - Shapes of Emotions: Multimodal Emotion Recognition in Conversations via
Emotion Shifts [2.443125107575822]
会話における感情認識(ERC)は重要かつ活発な研究課題である。
最近の研究は、ERCタスクに複数のモダリティを使用することの利点を示している。
マルチモーダルERCモデルを提案し,感情シフト成分で拡張する。
論文 参考訳(メタデータ) (2021-12-03T14:39:04Z) - Discovering Emotion and Reasoning its Flip in Multi-Party Conversations
using Masked Memory Network and Transformer [16.224961520924115]
感情フリップ推論(EFR)の新たな課題について紹介する。
EFRは、ある時点で感情状態が反転した過去の発話を特定することを目的としている。
後者のタスクに対して,前者およびトランスフォーマーベースのネットワークに対処するためのマスクメモリネットワークを提案する。
論文 参考訳(メタデータ) (2021-03-23T07:42:09Z) - Modality-Transferable Emotion Embeddings for Low-Resource Multimodal
Emotion Recognition [55.44502358463217]
本稿では、上記の問題に対処するため、感情を埋め込んだモダリティ変換可能なモデルを提案する。
我々のモデルは感情カテゴリーのほとんどで最先端のパフォーマンスを達成する。
私たちのモデルは、目に見えない感情に対するゼロショットと少数ショットのシナリオにおいて、既存のベースラインよりも優れています。
論文 参考訳(メタデータ) (2020-09-21T06:10:39Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。