STATISTICS

Viewed3

Downloads4

Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating
[1]LI Lihong,LI Zhixun,LIU Weiwei,et al.Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating[J].Journal of Zhengzhou University (Engineering Science),2026,47(4):9-16.[doi:10.13705/j.issn.1671-6833.2026.04.002]
Copy
References:
[1] Chandrasekaran G, Nguyen T N, Hemanth D J. Multimodal sentimental analysis for social media applications: a comprehensive review[J]. WIREs Data Mining and Knowledge Discovery, 2021, 11(5): e1415.
[2] Lyu Xueqiang, Tian Chi, Zhang Le, et al. Multimodal sentiment analysis model integrating multi‑features and attention mechanism[J]. Data Analysis and Knowledge Discovery, 2024, 8(5): 91‑101.[吕学强, 田驰, 张乐, 等. 融合多特征和注意力机制的多模态情感分析模型[J]. 数据分析与知识发现, 2024, 8(5): 91‑101.]
[3] Morency L P, Mihalcea R, Doshi P. Towards multimodal sentiment analysis: harvesting opinions from the web[C]//Proceedings of the 13th International Conference on Multimodal Interfaces. New York: ACM, 2011: 169‑176.
[4] Wang Yifeng, He Jiahao, Wang Di, et al. Multimodal transformer with adaptive modality weighting for multimodal sentiment analysis[J]. Neurocomputing, 2024, 572: 127181.
[5] Liu Zhizhong, Zhou Bin, Chu Dianhui, et al. Modality translation‑based multimodal sentiment analysis under uncertain missing modalities[J]. Information Fusion, 2024, 101: 101973.
[6] Kim W, Son B, Kim I. ViLT: vision‑and‑language transformer without convolution or region supervision[PP/OL]. V2. arXiv (2021‑06‑10)[2025‑12‑09]. https://arxiv.org/abs/2102.03334v2.
[7] Liu Zijun, Cai Li, Yang Wenjie, et al. Sentiment analysis based on text information enhancement and multimodal feature fusion[J]. Pattern Recognition, 2024, 156: 110847.
[8] Dou Ziyi, Xu Yichong, Gan Zhe, et al. An empirical study of training end‑to‑end vision‑and‑language transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 18145‑18155.
[9] Zeng Ying, Yan Wenjun, Mai Sijie, et al. Disentanglement Translation Network for multimodal sentiment analysis[J]. Information Fusion, 2024, 102: 102031.
[10] Chen Yan, Lai Yubin, Xiao Ao, et al. Multimodal sentiment analysis model based on CLIP and cross‑attention[J]. Journal of Zhengzhou University (Engineering Science), 2024, 45(2): 42‑50.[陈燕, 赖宇斌, 肖澳, 等. 基于CLIP和交叉注意力的多模态情感分析模型[J]. 郑州大学学报(工学版), 2024, 45(2): 42‑50.]
[11] Khan M, Tran P N, Pham N T, et al. MemoCMT: multimodal emotion recognition using cross‑modal transformer‑based feature fusion[J]. Scientific Reports, 2025, 15: 5473.
[12] Baevski A, Hsu W N, Xu Q T, et al. Data2vec: a general framework for self‑supervised learning in speech, vision and language[PP/OL]. V3. arXiv (2022‑10‑25)[2025‑12‑09]. https://arxiv.org/abs/2202.03555.
[13] Cui Yiming, Che Wanxiang, Liu Ting, et al. Pre‑training with whole word masking for Chinese BERT[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3504‑3514.
[14] Zhang Hang, Wu Chongruo, Zhang Zhongyue, et al. ResNetSt: split‑attention networks[C]//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Piscataway: IEEE, 2022: 2735‑2745.
[15] Baevski A, Zhou H, Mohamed A, et al. Wav2vec2.0: a framework for self‑supervised learning of speech representations[PP/OL]. arXiv (2020‑10‑22)[2025‑12‑09]. https://arxiv.org/abs/2006.11477.
[16] Huiyan A, Huang J X. STCA: utilizing a spatio‑temporal cross‑attention network for enhancing video person re‑identification[J]. Image and Vision Computing, 2022, 123: 104474.
[17] Bagher Zadeh A, Liang P P, Poria S, et al. Multimodal language analysis in the wild: CMU‑MOSEI dataset and interpretable dynamic fusion graph[C]//The 56th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2018: 2236‑2246.
[18] Sun Hao, Liu Jiaqing, Chen Y W, et al. Modality‑invariant and temporal representation learning for multimodal sentiment classification[J]. Information Fusion, 2023, 91: 504‑514.
[19] Golagana V, Row S V, Rao P S. Adaptive multimodal sentiment analysis: improving fusion accuracy with dynamic attention for missing modality[J]. Journal of Electrical Systems, 2024, 20(S1): 134‑147.
[20] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[PP/OL]. V7. arXiv (2023‑08‑02)[2025‑12‑09]. https://arxiv.org/abs/1706.03762.
[21] Xiao Luwei, Wu Xingjiao, Yang Shuwen, et al. Cross‑modal fine‑grained alignment and fusion network for multimodal aspect‑based sentiment analysis[J]. Information Processing & Management, 2023, 60(6): 103508.
[22] Zhao Fei, Zhang Chengcui, Geng Baocheng. Deep multimodal data fusion[J]. ACM Computing Surveys, 2024, 56(9): 1‑36.
[23] Lok E J. Toronto emotional speech set (TESS)[DS/OL]. [2025‑12‑09]. https://www.kaggle.com/datasets/ejlok1/toronto‑emotional‑speech‑set‑tess.
[24] Diem L, Zaharieva M. Video content representation using recurring regions detection[J]. Lecture Notes in Computer Science, 2016, 9516: 16‑28.
[25] Zhang Haoyu, Wang Yu, Yin Guanghao, et al. Learning language‑guided adaptive hyper‑modality representation for multimodal sentiment analysis[PP/OL]. V2. arXiv (2023‑12‑14)[2025‑12‑09]. https://arxiv.org/abs/2310.05804.
[26] Zadeh A, Liang P P, Mazumder N, et al. Memory fusion network for multi‑view sequential learning[PP/OL]. V1. arXiv (2018‑02‑03)[2025‑12‑09]. https://arxiv.org/abs/1802.00927.
[27] Tsai Y H, Bai Shaojie, Liang P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2019: 6558‑6569.
[28] Wang Di, Guo Xutong, Tian Yumin, et al. TETFN: a text enhanced transformer fusion network for multimodal sentiment analysis[J]. Pattern Recognition, 2023, 136: 109259.
[29] Zadeh A, Chen Minghai, Poria S, et al. Tensor fusion network for multimodal sentiment analysis[PP/OL]. V1. arXiv (2017‑07‑23)[2025‑12‑09]. https://arxiv.org/abs/1707.07250.
[30] Liu Zhun, Shen Ying, Lakshminarasimhan V B, et al. Efficient low‑rank multimodal fusion with modality‑specific factors[PP/OL]. V1. arXiv (2018‑05‑31)[2025‑12‑09]. https://arxiv.org/abs/1806.00064.
[31] Hu Jingwen, Liu Yuchen, Zhao Jinming, et al. MMGCN: multimodal fusion via deep graph convolution network for emotion recognition in conversation[PP/OL]. V1. arXiv (2021‑07‑14)[2025‑12‑09]. https://arxiv.org/abs/2107.06779.
[32] Ghosal D, Majumder N, Poria S, et al. DialogueGCN: a graph convolutional neural network for emotion recognition in conversation[PP/OL]. V1. arXiv (2019‑08‑30)[2025‑12‑09]. https://arxiv.org/abs/1908.11540.
[33] Hazarika D, Zimmermann R, Poria S. MISA: modality‑invariant and modality‑specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York: ACM, 2020: 1122‑1131.
[34] Xing Tao, Dou Yutao, Chen Xianliang, et al. An adaptive multi‑graph neural network with multimodal feature fusion learning for MDD detection[J]. Scientific Reports, 2024, 14: 28400.
Similar References:
Memo

-

Last Update: 2026-09-22
Copyright © 1980 Editorial Board of Journal of Zhengzhou University (Engineering Science)
Email: gxb@zzu.edu.cn ;Tel: 0371-67781276,0371-67781277
Address: No.100 Science Avenue,100,Zhengzhou 450001,China