[1]李丽红,李志勋,刘威伟,等.跨模态时空注意力与上下文门控的情感分析[J].郑州大学学报(工学版),2026,47(4):9-16.[doi:10.13705/j.issn.1671-6833.2026.04.002]
 LI Lihong,LI Zhixun,LIU Weiwei,et al.Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating[J].Journal of Zhengzhou University (Engineering Science),2026,47(4):9-16.[doi:10.13705/j.issn.1671-6833.2026.04.002]
点击复制

跨模态时空注意力与上下文门控的情感分析()
分享到:

《郑州大学学报(工学版)》[ISSN:1671-6833/CN:41-1339/T]

卷:
47
期数:
2026年4期
页码:
9-16
栏目:
出版日期:
2026-07-10

文章信息/Info

Title:
Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating
文章编号:
1671-6833(2026)04-0009-08
作者:
李丽红1,2, 李志勋1,2, 刘威伟1,2, 秦肖阳1,2
1. 华北理工大学 理学院,河北 唐山 063210;2. 河北省数据科学与应用重点实验室( 华北理工大学) ,河北 唐山 063210
Author(s):
LI Lihong1,2, LI Zhixun1,2, LIU Weiwei1,2, QIN Xiaoyang1,2
1. College of Science, North China University of Science and Technology, Tangshan 063210, China; 2. Hebei Province Key Laboratory of Data Science and Application (North China University of Science and Technology) , Tangshan 063210, China
关键词:
多模态情感分析 时空注意力 上下文门控 Transformer 跨模态融合 跨模态交互
Keywords:
multimodal sentiment analysis spatio-temporal attention contextual gating Transformer cross-modal fusion cross-modal interaction
分类号:
TP391. 1TN912. 3
DOI:
10.13705/j.issn.1671-6833.2026.04.002
文献标志码:
A
摘要:
多模态情感分析因模态异质导致的交互不一致性、语言场景复杂性及静态跨模态注意力难以捕捉多模态数据时序动态性,限制了深层模态关联挖掘与情感分类性能。针对以上难题,提出一种多模态情感分析框架,引入跨模态时空注意力(CM‑STA)捕获文本、图像与音频的时空依赖,增强跨模态交互;上下文门控(CG)动态筛选情感表达强相关的特征,突出关键情感信息;Transformer跨模态融合交互(TCMFI)通过多头自注意力与双线性池化实现深层跨模态融合,提升融合效率。所提模型在公开数据集TESS(音频)和MVSA‑Multiple(文本、图像)上的实验准确率为81.45%、F1分数为80.84%、AUROC为96.40%,较最佳基线模型MISA分别提升0.95、0.24和7.91个百分点;计算复杂度的实验结果显示,所提模型占用GPU内存7.8 GB,GPU利用率98%。所提模型以低空间复杂度和高GPU利用率实现高效融合,性能优于对比基线模型。实验结果验证了所提模型在复杂多模态情感分析场景中具有优异的性能与鲁棒性。
Abstract:
In multimodal sentiment analysis, it is difficult to capture the temporal dynamics of multimodal data by interaction inconsistencies due to modality heterogeneity, the complexity of linguistic scenarios, and the inability of static cross‑modal attention, which limits deep modality correlation mining and sentiment classification performance. To address these challenges, a multimodal sentiment analysis framework was proposed, incorporating cross‑modal spatio‑temporal attention (CM‑STA) to capture spatio‑temporal dependencies among text, image, and audio, enhancing cross‑modal interactions. Contextual gating (CG) was used to dynamically filter features strongly correlated with emotional expressions, emphasizing key sentiment information. A Transformer cross‑modal fusion interaction (TCMFI) was used to leverage multi‑head self‑attention and bilinear pooling for efficient deep cross‑modal fusion. Experiments on the TESS (audio) and MVSA‑Multiple (text, image) datasets yielded an accuracy of 81.45%, an F1 score of 80.84%, and an AUROC of 96.40%, outperforming the best baseline model MISA by 0.95, 0.24, and 7.91 percentage points, respectively. Computational complexity analysis revealed that the proposed model occupied 7.8 GB of GPU memory with a 98% GPU utilization rate, achieving efficient fusion with low spatial complexity and high GPU utilization, surpassing baseline models in performance. These results demonstrated the superior performance and robust effectiveness of the proposed model in complex multimodal sentiment analysis scenarios.

参考文献/References:

[1] Chandrasekaran G, Nguyen T N, Hemanth D J. Multimodal sentimental analysis for social media applications: a comprehensive review[J]. WIREs Data Mining and Knowledge Discovery, 2021, 11(5): e1415.
[2] Lyu Xueqiang, Tian Chi, Zhang Le, et al. Multimodal sentiment analysis model integrating multi‑features and attention mechanism[J]. Data Analysis and Knowledge Discovery, 2024, 8(5): 91‑101.[吕学强, 田驰, 张乐, 等. 融合多特征和注意力机制的多模态情感分析模型[J]. 数据分析与知识发现, 2024, 8(5): 91‑101.]
[3] Morency L P, Mihalcea R, Doshi P. Towards multimodal sentiment analysis: harvesting opinions from the web[C]//Proceedings of the 13th International Conference on Multimodal Interfaces. New York: ACM, 2011: 169‑176.
[4] Wang Yifeng, He Jiahao, Wang Di, et al. Multimodal transformer with adaptive modality weighting for multimodal sentiment analysis[J]. Neurocomputing, 2024, 572: 127181.
[5] Liu Zhizhong, Zhou Bin, Chu Dianhui, et al. Modality translation‑based multimodal sentiment analysis under uncertain missing modalities[J]. Information Fusion, 2024, 101: 101973.
[6] Kim W, Son B, Kim I. ViLT: vision‑and‑language transformer without convolution or region supervision[PP/OL]. V2. arXiv (2021‑06‑10)[2025‑12‑09]. https://arxiv.org/abs/2102.03334v2.
[7] Liu Zijun, Cai Li, Yang Wenjie, et al. Sentiment analysis based on text information enhancement and multimodal feature fusion[J]. Pattern Recognition, 2024, 156: 110847.
[8] Dou Ziyi, Xu Yichong, Gan Zhe, et al. An empirical study of training end‑to‑end vision‑and‑language transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 18145‑18155.
[9] Zeng Ying, Yan Wenjun, Mai Sijie, et al. Disentanglement Translation Network for multimodal sentiment analysis[J]. Information Fusion, 2024, 102: 102031.
[10] Chen Yan, Lai Yubin, Xiao Ao, et al. Multimodal sentiment analysis model based on CLIP and cross‑attention[J]. Journal of Zhengzhou University (Engineering Science), 2024, 45(2): 42‑50.[陈燕, 赖宇斌, 肖澳, 等. 基于CLIP和交叉注意力的多模态情感分析模型[J]. 郑州大学学报(工学版), 2024, 45(2): 42‑50.]
[11] Khan M, Tran P N, Pham N T, et al. MemoCMT: multimodal emotion recognition using cross‑modal transformer‑based feature fusion[J]. Scientific Reports, 2025, 15: 5473.
[12] Baevski A, Hsu W N, Xu Q T, et al. Data2vec: a general framework for self‑supervised learning in speech, vision and language[PP/OL]. V3. arXiv (2022‑10‑25)[2025‑12‑09]. https://arxiv.org/abs/2202.03555.
[13] Cui Yiming, Che Wanxiang, Liu Ting, et al. Pre‑training with whole word masking for Chinese BERT[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3504‑3514.
[14] Zhang Hang, Wu Chongruo, Zhang Zhongyue, et al. ResNetSt: split‑attention networks[C]//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Piscataway: IEEE, 2022: 2735‑2745.
[15] Baevski A, Zhou H, Mohamed A, et al. Wav2vec2.0: a framework for self‑supervised learning of speech representations[PP/OL]. arXiv (2020‑10‑22)[2025‑12‑09]. https://arxiv.org/abs/2006.11477.
[16] Huiyan A, Huang J X. STCA: utilizing a spatio‑temporal cross‑attention network for enhancing video person re‑identification[J]. Image and Vision Computing, 2022, 123: 104474.
[17] Bagher Zadeh A, Liang P P, Poria S, et al. Multimodal language analysis in the wild: CMU‑MOSEI dataset and interpretable dynamic fusion graph[C]//The 56th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2018: 2236‑2246.
[18] Sun Hao, Liu Jiaqing, Chen Y W, et al. Modality‑invariant and temporal representation learning for multimodal sentiment classification[J]. Information Fusion, 2023, 91: 504‑514.
[19] Golagana V, Row S V, Rao P S. Adaptive multimodal sentiment analysis: improving fusion accuracy with dynamic attention for missing modality[J]. Journal of Electrical Systems, 2024, 20(S1): 134‑147.
[20] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[PP/OL]. V7. arXiv (2023‑08‑02)[2025‑12‑09]. https://arxiv.org/abs/1706.03762.
[21] Xiao Luwei, Wu Xingjiao, Yang Shuwen, et al. Cross‑modal fine‑grained alignment and fusion network for multimodal aspect‑based sentiment analysis[J]. Information Processing & Management, 2023, 60(6): 103508.
[22] Zhao Fei, Zhang Chengcui, Geng Baocheng. Deep multimodal data fusion[J]. ACM Computing Surveys, 2024, 56(9): 1‑36.
[23] Lok E J. Toronto emotional speech set (TESS)[DS/OL]. [2025‑12‑09]. https://www.kaggle.com/datasets/ejlok1/toronto‑emotional‑speech‑set‑tess.
[24] Diem L, Zaharieva M. Video content representation using recurring regions detection[J]. Lecture Notes in Computer Science, 2016, 9516: 16‑28.
[25] Zhang Haoyu, Wang Yu, Yin Guanghao, et al. Learning language‑guided adaptive hyper‑modality representation for multimodal sentiment analysis[PP/OL]. V2. arXiv (2023‑12‑14)[2025‑12‑09]. https://arxiv.org/abs/2310.05804.
[26] Zadeh A, Liang P P, Mazumder N, et al. Memory fusion network for multi‑view sequential learning[PP/OL]. V1. arXiv (2018‑02‑03)[2025‑12‑09]. https://arxiv.org/abs/1802.00927.
[27] Tsai Y H, Bai Shaojie, Liang P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2019: 6558‑6569.
[28] Wang Di, Guo Xutong, Tian Yumin, et al. TETFN: a text enhanced transformer fusion network for multimodal sentiment analysis[J]. Pattern Recognition, 2023, 136: 109259.
[29] Zadeh A, Chen Minghai, Poria S, et al. Tensor fusion network for multimodal sentiment analysis[PP/OL]. V1. arXiv (2017‑07‑23)[2025‑12‑09]. https://arxiv.org/abs/1707.07250.
[30] Liu Zhun, Shen Ying, Lakshminarasimhan V B, et al. Efficient low‑rank multimodal fusion with modality‑specific factors[PP/OL]. V1. arXiv (2018‑05‑31)[2025‑12‑09]. https://arxiv.org/abs/1806.00064.
[31] Hu Jingwen, Liu Yuchen, Zhao Jinming, et al. MMGCN: multimodal fusion via deep graph convolution network for emotion recognition in conversation[PP/OL]. V1. arXiv (2021‑07‑14)[2025‑12‑09]. https://arxiv.org/abs/2107.06779.
[32] Ghosal D, Majumder N, Poria S, et al. DialogueGCN: a graph convolutional neural network for emotion recognition in conversation[PP/OL]. V1. arXiv (2019‑08‑30)[2025‑12‑09]. https://arxiv.org/abs/1908.11540.
[33] Hazarika D, Zimmermann R, Poria S. MISA: modality‑invariant and modality‑specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York: ACM, 2020: 1122‑1131.
[34] Xing Tao, Dou Yutao, Chen Xianliang, et al. An adaptive multi‑graph neural network with multimodal feature fusion learning for MDD detection[J]. Scientific Reports, 2024, 14: 28400.

更新日期/Last Update: 2026-09-22