2026 Volume 47 Issue Issue 4
[0]
Abstract:
CHEN Enqing, LI Jiahui, GUO Xin
Abstract: To address the problems of incomplete motion information caused by occlusion or missing joints in skeleton‑based action recognition, as well as the limited generalization ability of models with few‑label conditions, a skeleton‑based action recognition method DCMAE was proposed, which integrated a diffusion model with a cross‑attention mechanism. Within a self‑supervised learning framework, a spatio‑temporal masking strategy was adopted, where the diffusion model learned the global distribution characteristics of motion sequences during the denoising process to improve classification accuracy under data‑missing conditions. In the decoding stage, the cross‑attention mechanism introduced encoder features to achieve spatio‑temporal information interaction and guidance, thereby enhancing the model’s generalization ability in few‑label conditions. Experiments conducted on the NTU RGB+D 60 and NTU RGB+D 120 datasets showed that the proposed method could achieve accuracy improvements of up to 14.9 percentage points and 3.0 percentage points, respectively, over SkeletonMAE with data‑missing conditions and few‑label conditions. The proposed method effectively enhanced the robustness of skeleton‑based action recognition models to data‑missing and few‑label data, providing a new perspective for self‑supervised action recognition research.
LI Lihong1,2, LI Zhixun1,2, LIU Weiwei1,2, QIN Xiaoyang1,2
Abstract: In multimodal sentiment analysis, it is difficult to capture the temporal dynamics of multimodal data by interaction inconsistencies due to modality heterogeneity, the complexity of linguistic scenarios, and the inability of static cross‑modal attention, which limits deep modality correlation mining and sentiment classification performance. To address these challenges, a multimodal sentiment analysis framework was proposed, incorporating cross‑modal spatio‑temporal attention (CM‑STA) to capture spatio‑temporal dependencies among text, image, and audio, enhancing cross‑modal interactions. Contextual gating (CG) was used to dynamically filter features strongly correlated with emotional expressions, emphasizing key sentiment information. A Transformer cross‑modal fusion interaction (TCMFI) was used to leverage multi‑head self‑attention and bilinear pooling for efficient deep cross‑modal fusion. Experiments on the TESS (audio) and MVSA‑Multiple (text, image) datasets yielded an accuracy of 81.45%, an F1 score of 80.84%, and an AUROC of 96.40%, outperforming the best baseline model MISA by 0.95, 0.24, and 7.91 percentage points, respectively. Computational complexity analysis revealed that the proposed model occupied 7.8 GB of GPU memory with a 98% GPU utilization rate, achieving efficient fusion with low spatial complexity and high GPU utilization, surpassing baseline models in performance. These results demonstrated the superior performance and robust effectiveness of the proposed model in complex multimodal sentiment analysis scenarios.
GONG Qiuming1, LI Shunwen1, HUANG Liu1, WANG Ju2,3, CAO Zixiang1, MA Hongsu2,3
Abstract:
To address the limitations of existing surrounding rock mass identification methods based on TBM vibration signals in terms of feature extraction effectiveness and engineering adaptability,a novel surrounding rock mass perception method was proposed by integrating wavelet scattering network (WSN) and long short-term memory network ( LSTM) using TBM cutterhead vibration data. Firstly, relying on the spiral ramp project of the Beishan Underground Laboratory, a vibration monitoring system was mounted on the TBM cutterhead to acquire vibration signals during the TBM tunneling process. Then, a rock mass sensing database based on cutterhead vibration was established through a series of data preprocessing procedures, including stable tunneling segment extraction, noise reduction, and signal segmentation, combined with the matching of geological information along the tunnel alignment.
Secondly, the WSN was employed to perform multi-scale temporal feature extraction from the preprocessed vibration signals, so as to enhance the feature representation capability and noise robustness. On this basis, a WSN-LSTM surrounding rock mass perception model was constructed by leveraging the inherent superiority of the LSTM network in capturing the temporal dependencies. The results demonstrated that the proposed WSN-LSTM model achieved an accuracy of 93. 7% on the test set, which yielded a 5. 6 percentage points accuracy improvement compared with the wavelet scattering network-based support vector machine ( SVM) model, and outperformed shallow machine learning models ( random forest and LightGBM) based on amplitude-domain statistical feature extraction. These findings validated the superiority of WSN in feature extraction from TBM cutterhead vibration signals, as well as the necessity of capturing the temporal dependencies of cutterhead vibration features.

ZHANG Zhengqi1, HAN Yanzhi1, LEI Zhikun1,2, SHI Jierong3, YANG Xinhong3, YANG Mi3
Abstract:
To investigate the influencing factors and mechanisms of fume release from crumb rubber modified asphalt, based on the preparation of different crumb rubber modified asphalts, a self-developed fume generation and detection device was utilized, combined with three detection methods including gravimetric method, portable gas detector, and gas chromatography-mass spectrometry(GC-MS) to determine the concentrations of asphalt fumes and harmful components. The grey relational analysis was performed to evaluate the relationship between various factors and fume concentrations. Further characterization using four-component analysis, infrared spectroscopy, and fluorescence microscopy was conducted to explore how different factors affect the underlying mechanisms of fume release. The results showed that the base asphalt grade and additive type were the main factors affecting asphalt fume release. The four-component test showed that fumes from crumb rubber modified asphalt mainly came from the volatilization of light components. The infrared spectroscopy test indicated that the contents of aromatic hydrocarbons and alkanes were the main factors affecting fume and VOCs concentrations,while the total sulfur content in asphalt was the key factor controlling H2 S release. The fluorescence microscopy test further confirmed that a stable internal structure could suppress fume release from crumb rubber modified asphalt to a certain extent.

SUN Xiao1, WANG Xiangyang1, YANG Zhuanjia2, ZHANG Xinyu1
Abstract:
Aiming at the dispersion problem of underwater crack repair materials for concrete dams, an underwater non-dispersible grouting material containing diatomite was designed. Firstly, the mesh number and content of diatomite were determined by mercury intrusion test and single mixing test, and the influence of diatomite on the mechanical properties of grouting materials was analyzed by compressive strength and splitting tensile strength tests. Secondly, the effects of hydroxypropyl methyl cellulose flocculant (HPMC) and ordinary PCA®-typeⅠ polycarboxylic acid high performance water reducing agent on the flowability and anti-washout performance of grouting materials were further studied by cone flowability method, visual observation method, pH value method, and plunge test. Finally, based on the orthogonal test method, the specific ratio of underwater grouting repair materials was determined. The results showed that when the water-cement ratio was 0. 50, the diatomite had good compatibility with the slurry, and the addition of 2% ( mass fraction, the same below) 100 mesh diatomite could improve the compressive strength and splitting tensile strength of the grouting material. On this basis, the addition of 0. 6% HPMC could improve the anti-washout performance of the slurry, and the addition of 0. 10% water reducing agent could improve its flowability. The diatomite mesh of underwater grouting material was determined to be 100 mesh, the content was 2%, the content of HPMC was 0. 6%, and the content of PCA®-typeⅠ polycarboxylic acid high performance water reducing agent was 0. 10%.

WANG Jingyang1, XU Yongchao1, ZHANG Bo2, WANG Jue3, HUANG Min1
Abstract: Aiming at the problem that the existing road crack detection model cannot effectively balance the detection accuracy, computational complexity and detection speed and has poor practical application effect, a lightweight road crack detection model YOLO-CGVE ba<x>sed on improved YOLOv10n is proposed. Firstly, the coordinate attention (CA) module is used to replace the partial self-attention (PSA) module to better capture the local and global relationships in space and improve the capacity to extract features. Secondly, the computational complexity is reduced by using lightweight GSConv to replace some standard convolution structures. Then, the original C2f structure in the neck network is replaced by VoV-GSCSP, which allows for the efficient merging of feature maps from various stages and further minimizes computing complexity while maintaining accuracy. Finally, the ECIoU loss function is used to replace the original loss function to improve the detection box positioning accuracy and convergence speed. The experimental results on the public dataset RDD2022 show that compared with YOLOv10n, while keeping a high detection speed, the mAP@0.5 of YOLO-CGVE is improved by 2.4 percentage points, reaching 75.9%, and the parameters and the GFLOPs are decreased by 11.1% and 9.8%, respectively. YOLO-CGVE can better meet the application needs in environments with limited computing resources.
ZHANG Zhen1,2, CHUN Meijie1, TIAN Hongpeng2, LI Youhao3, HUANG Weitao3, ZHANG Junjie3
Abstract: In response to the issue of biased estimates affecting classification performance in imputation-based classification methods when dealing with missing data, an incomplete data evidence ensemble classification method based on adaptive subspace imputation was proposed. The proposed method utilized adaptive subspace imputation and dual evidence integration to enhance the model′s classification ability on incomplete datasets. Firstly, spectral clustering was used to dynamically partition the feature space into multiple subspaces, where missing value imputation based on neighbors was performed independently within each subspace. Secondly, a dual importance evaluation mechanism was designed, which calculated the difference in data distribution before and after imputation in the training set to assess global importance, and evaluated the local importance of classification results by assessing the classification capacity of the classification model on the test set samples′ neighbors in the training set. Finally, based on evidence theory, local and global importance were fused to enhance classification performance by leveraging the complementarity of information from different subspaces. Comparative experiments on standard datasets showed that the proposed method achieved improvements of up to 6. 23 percentage and 0. 82 percentage, respectively, in the ARI and AP metrics compared to suboptimal methods, validating the effectiveness and advancement of the proposed method.
ZHENG Hong, LUO Yujian, LING Kan, FAN Guisheng
Abstract:
To address the current challenges in relying solely on single-modal features of target proteins and neglecting network-scale features of biological networks, a drug-target affinity prediction (DTA) model based on multimodal cross-scale feature fusion was proposed. Target proteins both as sequences and graphs for feature extraction were presented, with semantic and topological features respectively, to enhance the target proteins′ feature representation. The strong affinity relationships between drugs and target proteins were analyzed to construct a heterogeneous graph network of drug-target interaction. A cross-scale feature fusion method was then used to effectively integrate the scale features of the heterogeneous graph network, then to enrich the feature representations of both target proteins and drug molecules. Experimental results on the DAVIS and KIBA datasets demonstrated that, compared with the more advanced model SISDTA, the proposed model achieved reductions in MSE by 0. 015 and 0. 003, respectively, and increases in CI by 0. 005 and 0. 004, respectively, improving the accuracy and stability of affinity prediction. It demonstrated the effectiveness of multimodal and cross-scale feature fusion in DTA prediction tasks.

LIU Jing1,2, JIANG Wenjie1, FENG Hailing3, ZHANG Haibin4, JI Haipeng2,3,5
Abstract: Aiming at the problem of the disconnection between domain knowledge and data‑driven models in traditional oxygen supply prediction methods in converter steelmaking process, a knowledge and data fusion driven oxygen supply prediction method for converter steelmaking was proposed. A three‑level knowledge fusion module was constructed, embedding metallurgical mechanisms into deep learning models. Secondly, a dual‑branch architecture was designed to collaboratively mine process characteristics and cross furnace production patterns. Finally, actual production data from a domestic steel plant was used for the experiment. The experimental results showed that compared with mainstream methods such as GBRBM‑DBN, HyGPR, Stacking, and BOA‑LGBM, the MAE and RMSE of oxygen supply with SPHC steel grade decreased by a maximum of 7.59% and 6.80% respectively, and the accuracy (relative error ±5%) reached 85.29%. With the HRB400E steel grade, the MAE and RMSE decreased by a maximum of 15.24% and 15.13% respectively, with an accuracy (relative error ±5%) of 87.91%, verifying the oxygen supply prediction ability of the proposed method.
ZHANG Guangchen1, LI Zhanfei1, HE Shuping2, XIA Yuanqing3
Abstract: For the classification problem of linearly inseparable datasets, a support vector machine (SVM) kernel function parameter optimization algorithm was proposed based on the sliding mode control (SMC) strategy by applying the SMC idea to the SVM kernel function parameter optimization process. By designing the error equation and sliding surface, the association between SVM classification objective function and SMC was established, and then the iteration update rules of kernel function parameters and cost function was derived. The algorithm improved classification performance while reducing the number of support vectors by dynamically adjusting the kernel parameters of SVM. In the experiment part, six UCI datasets, such as Iris and Heart disease, were used to verify the validity of the algorithm. The results showed that compared with traditional SVM, the proposed algorithm reduced the number of support vectors by 56.25% on the Iris dataset, and the test accuracy remained at 100%. Test accuracy increased by 13.58 percentage points on the Heart disease dataset. Furthermore, the proposed algorithm, compared with existing optimization algorithms, showed a higher classification accuracy on some datasets.
LI Suyue, LI Juan, ZHANG Yabin, WANG Anhong
Abstract: An active reconfigurable intelligent surface (ARIS)-aided rate‑splitting multiple access (RSMA) system in a multi‑antenna multi‑user scenario was investigated, and a transmission framework that integrated base station antenna selection with user side combining techniques was proposed. The statistical characteristics of cascaded user channels were derived based on selection combining (SC) and maximal ratio combining (MRC), then an outage probability analysis model was established, yielding analytical expressions for the outage probability in multi‑antenna multi‑user scenarios. Furthermore, to minimize the system outage probability, for both combining schemes, power allocation optimization problems were formulated and solved via a differential evolution (DE) algorithm. Simulation results demonstrated that deploying multiple antennas at the user side and adopting appropriate antenna selection strategies could significantly reduce the outage probability, thereby confirm the scalability and robustness of the proposed ARIS‑RSMA framework in complex communication environments.
JI Xinfang1, 2, JIA Jingwei1, 2, WANG Xiaofeng1, 2, CHENG Jinxin3, YAO Jiaxing1, 2
Abstract: Expensive multimodal optimization problems (EMMOPs) arise in engineering design frequently and are often characterized as multimodal properties and with extremely high evaluation costs. The progress and key techniques of surrogate‑assisted evolutionary algorithms (SAEAs) for such problems were systematically reviewed in the study. Firstly, typical surrogate models, including polynomial regression model and Gaussian process, were introduced, with emphasis on their characteristics and applicability in sample fitting, nonlinear representation, and uncertainty quantification. Then, the general framework of SAEAs was summarized, and the main design ideas of existing algorithms were outlined in terms of single‑surrogate and multi‑surrogate structures, global‑local collaborative search, and infill sampling strategies. Subsequently, according to the different characteristics of EMMOPs, typical EMMOPs, including single‑objective, multi‑objective, constrained, and high‑dimensional problems, were systematically categorized and reviewed, with particular attention to advances in mode identification, solution diversity preservation, and computational budget allocation. Furthermore, experimental comparisons of multiple mainstream SAEAs were conducted on ten benchmark test problems, and the performance differences among various algorithms were analyzed in terms of metrics such as global optimum solution and effective valley ratio. Meanwhile, engineering case studies, including ship structure optimization and synchronous machine design in ultra‑high‑voltage direct current transmission systems, were incorporated to illustrate the application potential of surrogate‑assisted evolutionary algorithms in complex engineering optimization. Finally, the key challenges faced by current research were summarized, and future development directions were discussed from the perspectives of adaptive surrogate model management, parallel execution and scheduling, as well as inter‑modal information sharing and transfer mechanisms.
MA Li1,2, LIU Wenzhe1, LI Yuhao1
Abstract: Abstract: To address the challenges of inefficient information fusion and noise interference in sequential recommendation, a novel method based on an adaptive bidirectional information flow was proposed. Built upon a dual‑path encoder architecture, a hierarchical history summarization module was integrated to distill long‑term user preferences, and dynamic frequency‑domain filtering was introduced to suppress data noise. The approach fully considered the dependency and interactivity between past and future information by employing an adaptive bidirectional information flow mechanism. This mechanism dynamically adjusted fusion weights via uncertainty perception, enabling a precise characterization of the evolution of user preferences. To validate its effectiveness, experiments were conducted on four public datasets including Beauty, Sports, Yelp, and ML1M. And a comparative analysis was performed against 10 mainstream methods. The experimental results demonstrated that the proposed method out‑performed the baseline models in three key metrics: NDCG, HR, and MRR. Compared to three leading baseline models of FMLP‑Rec, DualRec, and Oracle4Rec, the proposed method’s HR@20 reached 0.652 0 and 0.913 3 on the Beauty and Yelp datasets, which was 2.08 percentage points and 2.89 percentage points higher than their average performance, respectively. Furthermore, its NDCG@20 on the Beauty and Yelp datasets reached 0.394 4 and 0.564 5, outperforming the average of the three baselines by 2.67 percentage points and 2.72 percentage points, respectively.
ZHANG Jianhui1,2, XU Sijie1, ZENG Junjie1, WANG Ruimin3
Abstract: To address the problem that once discretely triggered mutation‑based moving target defense (MTD) strategies in digital twin network (DTN) could not continuously intercept malicious traffic during trigger intervals, which might result in protection gaps, a mutation‑service deception collaborative MTD method, termed MSD‑MTD was proposed. Building upon address and service port mutation, MSD‑MTD introduced a service deception mechanism to redirect suspicious traffic within mutation intervals, thereby enhancing continuous protection. Moreover, an intrusion detection approach based on cross‑node traffic alignment and feature selection was employed to perceive network states, and a deep Q‑network (DQN) was used to enable adaptive selection of MTD strategies. Comparative experiments were conducted on the Mininet‑WiFi platform using the CICIDS‑2017, CICIDS‑2018, and UNSW‑NB15 datasets, with performance benchmarked against two representative address‑mutation methods. The results showed that MSD‑MTD achieved average defense success rates of 93.36%, 88.20%, and 95.50% on the three datasets respectively, while the round‑trip time was mainly distributed within 0‑2 ms, indicating that the proposed method improved defense effectiveness while imposing only limited impact on network service latency.
YAN Hongcan1,2, ZHAO Yuting1, LI Sijia3, XIN Yuchi1
Abstract: The exponential growth of mobile trajectory data in location‑based services has significantly increased the risk of user privacy leakage. It is urgent and necessary to make effective privacy protection mechanisms. To enhance the utility of trajectory data while ensuring privacy protection, a trajectory privacy protection model named TCI‑BiGAN was constructed based on BiLSTM‑GAN. The Bayesian optimization method was used to perform adaptive parameter tuning for hierarchical density‑based spatial clustering of applications with noise (HDBSCAN), to improve data processing efficiency and reduce trajectory redundancy. BiLSTM was embedded into both the generator and discriminator of the generative adversarial network to efficiently extract spatiotemporal features and capture dependencies of trajectory data through its contextual feature extraction capability, thereby to enhance the similarity between generated and real trajectories. A multivariate discrete hidden Markov model was applied for trajectory interpolation, to increase data completeness and utility. On the Foursquare NYC and T‑Drive real‑world datasets, the user trajectory linkage accuracy was reduced to 0.243 and 0.198 respectively, and the average Hausdorff distance between generated and real trajectories was decreased to 0.013 and 0.019 respectively.
ZHANG Shufen1,2,3, LI Tao1,2,3, ZHANG Zhenbo1,2,3, ZHONG Qi1,2,3, JING Zhongrui1,2,3
Abstract: To address the issue that existing defense schemes in federated learning tend to over‑prune benign models during filtering, a robust aggregation algorithm defending against Byzantine attacks in federated learning (FLDBA) was proposed. HDBSCAN density‑based clustering was employed to group models, to identify the benign cluster, and the most representative model in terms of direction was selected as the trusted reference model. Using the trusted model as a benchmark, cosine similarity was utilized to screen potentially misclassified benign models within clusters, thereby to correct misjudgments. Additionally, a reputation mechanism was established to dynamically evaluate models’ historical behaviors, to mitigate the impact of missed detections. For models with high reputation, adaptive magnitude scaling was applied, and differential aggregation weights were assigned based on update quality to further enhance aggregation performance. Experimental results demonstrated that when defending against sign‑flipping attacks, FLDBA achieved an accuracy improvement of 0.18 percentage points to 5.13 percentage points compared to FLRAM, FLAME, RFLPA, FLTrust, and Krum, while reducing the attack success rate by 40.52 percentage points to 61.39 percentage points, exhibiting superior robustness.
HAN Jihui1, SHI Yupeng1, HUANG Ziqi2, ZHANG Anlin3, HUANG Daoying1
Abstract: To address the degradation of node representations in graph neural networks with complex perturbation environments, a structure‑feature collaborative defense graph neural network named SFCoRobustGNN was proposed. Structurally, a sparse attention mechanism that integrated structure priors to dynamically suppress anomalous edges was introduced. Feature‑wise, a channel gating mechanism was combined with a nonlinear feature mixing module (FeatureMixPro) to enhance the model’s adaptability to feature perturbations. A collaborative dual‑pathway defense was achieved through adversarial training and a multi‑objective optimization strategy. Experiments on multiple benchmark datasets, including Cora and Citeseer, demonstrated that the proposed method outperformed most mainstream baseline methods with various intensities of structure perturbations (5%‑40%) and feature attacks (\varepsilon = 0.01‑0.10), showing significant improvement in node classification accuracy. On the large‑scale ogbn‑products dataset, it maintained an accuracy of 71.82% even with a 20% MetaAttack structure perturbation, demonstrating its strong scalability. Ablation studies validated the effectiveness and synergistic effects of each module. The proposed method effectively mitigated performance degradation with complex perturbations and exhibited excellent generalization.
Copyright © 1980 Editorial Board of Journal of Zhengzhou University (Engineering Science)
Email: gxb@zzu.edu.cn ;Tel: 0371-67781276,0371-67781277
Address: No.100 Science Avenue,100,Zhengzhou 450001,China