Gated Mechanism for Attention Based Multimodal Sentiment Analysis
- URL: http://arxiv.org/abs/2003.01043v1
- Date: Fri, 21 Feb 2020 06:58:03 GMT
- Title: Gated Mechanism for Attention Based Multimodal Sentiment Analysis
- Authors: Ayush Kumar, Jithendra Vepa
- Abstract summary: Multimodal sentiment analysis has recently gained popularity because of its relevance to social media posts, customer service calls and video blogs.
In this paper, we address three aspects of multimodal sentiment analysis; 1. Cross modal interaction learning, i.e. how multiple modalities contribute to the sentiment.
We perform experiments on two benchmark datasets, CMU Multimodal Opinion level Sentiment Intensity (CMU-MOSI) and CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus.
- Score: 7.07652817535224
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Multimodal sentiment analysis has recently gained popularity because of its
relevance to social media posts, customer service calls and video blogs. In
this paper, we address three aspects of multimodal sentiment analysis; 1. Cross
modal interaction learning, i.e. how multiple modalities contribute to the
sentiment, 2. Learning long-term dependencies in multimodal interactions and 3.
Fusion of unimodal and cross modal cues. Out of these three, we find that
learning cross modal interactions is beneficial for this problem. We perform
experiments on two benchmark datasets, CMU Multimodal Opinion level Sentiment
Intensity (CMU-MOSI) and CMU Multimodal Opinion Sentiment and Emotion Intensity
(CMU-MOSEI) corpus. Our approach on both these tasks yields accuracies of 83.9%
and 81.1% respectively, which is 1.6% and 1.34% absolute improvement over
current state-of-the-art.
Related papers
- MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model [57.89395815934156]
Multi-Turn Contrastive Learning (MuCo) is a dialogue-inspired framework that revisits this process.<n>Experiments exhibit MuCo with a newly curated 5M multimodal multi-turn dataset (M3T)
arXiv Detail & Related papers (2026-02-06T05:18:33Z) - MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition [8.732416479560605]
This paper proposes Multimodal Cross-Attention Network and Contrastive Learning (MCN-CL) for multimodal emotion recognition.<n>It uses a triple query mechanism and hard negative mining strategy to remove feature redundancy while preserving important emotional cues.<n>Experiment results on the IEMOCAP and MELD datasets show that our proposed method outperforms state-of-the-art approaches.
arXiv Detail & Related papers (2025-11-14T02:13:31Z) - Hierarchical Adaptive Expert for Multimodal Sentiment Analysis [5.755715236558973]
Multimodal sentiment analysis has emerged as a critical tool for understanding human emotions across diverse communication channels.
We propose the Hierarchical Adaptive Expert for Multimodal Sentiment Analysis (HAEMSA), a novel framework that combines evolutionary optimization, cross-modal knowledge transfer, and multi-task learning.
Extensive experiments demonstrate HAEMSA's superior performance across multiple benchmark datasets.
arXiv Detail & Related papers (2025-03-25T09:52:08Z) - MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark [77.93283927871758]
This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning benchmark.
MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities.
arXiv Detail & Related papers (2024-09-04T15:31:26Z) - CorMulT: A Semi-supervised Modality Correlation-aware Multimodal Transformer for Sentiment Analysis [2.3522423517057143]
We propose a two-stage semi-supervised model termed Correlation-aware Multimodal Transformer (CorMulT)
At the pre-training stage, a modality correlation contrastive learning module is designed to efficiently learn modality correlation coefficients between different modalities.
At the prediction stage, the learned correlation coefficients are fused with modality representations to make the sentiment prediction.
arXiv Detail & Related papers (2024-07-09T17:07:29Z) - Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism [41.16574023720132]
This study introduces a novel architecture, named Multimodal Stable Fusion with Gated Cross-Attention (MSGCA), designed to robustly integrate multimodal input for stock movement prediction.
MSGCA framework consists of three integral components: (1) a trimodal encoding module, responsible for processing indicator sequences, dynamic documents, and a relational graph, and standardizing their feature representations; (2) a cross-feature fusion module, where primary and consistent features guide the multimodal fusion of the three modalities via a pair of gated cross-attention networks; and (3) a prediction module, which refines the fused features through temporal and dimensional reduction to execute precise
arXiv Detail & Related papers (2024-06-06T03:13:34Z) - Deep Equilibrium Multimodal Fusion [88.04713412107947]
Multimodal fusion integrates the complementary information present in multiple modalities and has gained much attention recently.
We propose a novel deep equilibrium (DEQ) method towards multimodal fusion via seeking a fixed point of the dynamic multimodal fusion process.
Experiments on BRCA, MM-IMDB, CMU-MOSI, SUN RGB-D, and VQA-v2 demonstrate the superiority of our DEQ fusion.
arXiv Detail & Related papers (2023-06-29T03:02:20Z) - Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications [90.6849884683226]
We study the challenge of interaction quantification in a semi-supervised setting with only labeled unimodal data.
Using a precise information-theoretic definition of interactions, our key contribution is the derivation of lower and upper bounds.
We show how these theoretical results can be used to estimate multimodal model performance, guide data collection, and select appropriate multimodal models for various tasks.
arXiv Detail & Related papers (2023-06-07T15:44:53Z) - Cross-Attention is Not Enough: Incongruity-Aware Dynamic Hierarchical
Fusion for Multimodal Affect Recognition [69.32305810128994]
Incongruity between modalities poses a challenge for multimodal fusion, especially in affect recognition.
We propose the Hierarchical Crossmodal Transformer with Dynamic Modality Gating (HCT-DMG), a lightweight incongruity-aware model.
HCT-DMG: 1) outperforms previous multimodal models with a reduced size of approximately 0.8M parameters; 2) recognizes hard samples where incongruity makes affect recognition difficult; 3) mitigates the incongruity at the latent level in crossmodal attention.
arXiv Detail & Related papers (2023-05-23T01:24:15Z) - Multi-channel Attentive Graph Convolutional Network With Sentiment
Fusion For Multimodal Sentiment Analysis [10.625579004828733]
This paper proposes a Multi-channel Attentive Graph Convolutional Network (MAGCN)
It consists of two main components: cross-modality interactive learning and sentimental feature fusion.
Experiments are conducted on three widely-used datasets.
arXiv Detail & Related papers (2022-01-25T12:38:33Z) - Channel Exchanging Networks for Multimodal and Multitask Dense Image
Prediction [125.18248926508045]
We propose Channel-Exchanging-Network (CEN) which is self-adaptive, parameter-free, and more importantly, applicable for both multimodal fusion and multitask learning.
CEN dynamically exchanges channels betweenworks of different modalities.
For the application of dense image prediction, the validity of CEN is tested by four different scenarios.
arXiv Detail & Related papers (2021-12-04T05:47:54Z) - Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal
Sentiment Analysis [96.46952672172021]
Bi-Bimodal Fusion Network (BBFN) is a novel end-to-end network that performs fusion on pairwise modality representations.
Model takes two bimodal pairs as input due to known information imbalance among modalities.
arXiv Detail & Related papers (2021-07-28T23:33:42Z) - Video Sentiment Analysis with Bimodal Information-augmented Multi-Head
Attention [7.997124140597719]
This study focuses on the sentiment analysis of videos containing time series data of multiple modalities.
The key problem is how to fuse these heterogeneous data.
Based on bimodal interaction, more important bimodal features are assigned larger weights.
arXiv Detail & Related papers (2021-03-03T12:30:11Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.