[1]LI Lihong,LI Zhixun,LIU Weiwei,et al.Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating[J].Journal of Zhengzhou University (Engineering Science),2026,47(4):9-16.[doi:10.13705/j.issn.1671-6833.2026.04.002]
Copy
Journal of Zhengzhou University (Engineering Science)[ISSN
1671-6833/CN
41-1339/T] Volume:
47
Number of periods:
2026 Issue 4
Page number:
9-16
Column:
Public date:
2026-07-10
- Title:
-
Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating
- Author(s):
-
LI Lihong1,2, LI Zhixun1,2, LIU Weiwei1,2, QIN Xiaoyang1,2
-
1. College of Science, North China University of Science and Technology, Tangshan 063210, China; 2. Hebei Province Key Laboratory of Data Science and Application (North China University of Science and Technology) , Tangshan 063210, China
-
- Keywords:
-
multimodal sentiment analysis; spatio-temporal attention; contextual gating; Transformer cross-modal fusion; cross-modal interaction
- CLC:
-
TP391. 1TN912. 3
- DOI:
-
10.13705/j.issn.1671-6833.2026.04.002
- Abstract:
-
In multimodal sentiment analysis, it is difficult to capture the temporal dynamics of multimodal data by interaction inconsistencies due to modality heterogeneity, the complexity of linguistic scenarios, and the inability of static cross‑modal attention, which limits deep modality correlation mining and sentiment classification performance. To address these challenges, a multimodal sentiment analysis framework was proposed, incorporating cross‑modal spatio‑temporal attention (CM‑STA) to capture spatio‑temporal dependencies among text, image, and audio, enhancing cross‑modal interactions. Contextual gating (CG) was used to dynamically filter features strongly correlated with emotional expressions, emphasizing key sentiment information. A Transformer cross‑modal fusion interaction (TCMFI) was used to leverage multi‑head self‑attention and bilinear pooling for efficient deep cross‑modal fusion. Experiments on the TESS (audio) and MVSA‑Multiple (text, image) datasets yielded an accuracy of 81.45%, an F1 score of 80.84%, and an AUROC of 96.40%, outperforming the best baseline model MISA by 0.95, 0.24, and 7.91 percentage points, respectively. Computational complexity analysis revealed that the proposed model occupied 7.8 GB of GPU memory with a 98% GPU utilization rate, achieving efficient fusion with low spatial complexity and high GPU utilization, surpassing baseline models in performance. These results demonstrated the superior performance and robust effectiveness of the proposed model in complex multimodal sentiment analysis scenarios.