[1]彭春燕,王 璇,陈杨博,等.基于图卷积网络的三维手部姿态估计[J].郑州大学学报(工学版),2026,47(5):9-16.[doi:10.13705/j.issn.1671-6833.2026.02.013]
 PENG Chunyan,WANG Xuan,CHEN Yangbo,et al.3D Hand Pose Estimation Based on Graph Convolution Network[J].Journal of Zhengzhou University (Engineering Science),2026,47(5):9-16.[doi:10.13705/j.issn.1671-6833.2026.02.013]
点击复制

基于图卷积网络的三维手部姿态估计()
分享到:

《郑州大学学报(工学版)》[ISSN:1671-6833/CN:41-1339/T]

卷:
47
期数:
2026年5期
页码:
9-16
栏目:
出版日期:
2026-09-10

文章信息/Info

Title:
3D Hand Pose Estimation Based on Graph Convolution Network
文章编号:
1671-6833(2026)05-0009-08
作者:
彭春燕1,2, 王 璇1,2, 陈杨博1,2, 何港波1,2
1. 青海师范大学 计算机学院,青海 西宁 810016;2. 青海师范大学 藏语智能全国重点实验室,青海 西宁 810016
Author(s):
PENG Chunyan1,2, WANG Xuan1,2, CHEN Yangbo1,2, HE Gangbo1,2
1. College of Computer, Qinghai Normal University, Xining 810016, China; 2. The State Key Laboratory of Tibetan Intelligence, Qinghai Normal University, Xining 810016, China
关键词:
三维手部姿态估计;  图卷积网络;  特征提取;  图核学习优化;  评估指标动态调整
Keywords:
3D hand pose estimation;  graph convolution networks;  feature extraction;  optimisation of graph kernel learning;  dynamic adjustment of assessment indicators
分类号:
TP391TP751
DOI:
10.13705/j.issn.1671-6833.2026.02.013
文献标志码:
A
摘要:
基于单张彩色图片的三维手部姿态估计由于手部存在遮挡、手部自相似性高等原因使预测结果存在误差大、手部结构不自然等问题。针对这些问题,首先,提出一个基于图卷积的三维手部姿态估计方法,使用Keypoint R‑CNN提取图像视觉特征和手部关键点二维位置信息,将特征信息输入到改进的自适应核图卷积模块(AK‑GraFormer)中;其次,引入带残差连接的AKGNN图核,自适应处理图数据以增强模型的特征学习与表达;最后,利用提出的评估指标监控动态训练策略以获得更优的估计结果。在HO‑3D_v3数据集与FreiHand数据集上进行实验,结果表明:在单张彩色图片手部三维姿态估计任务中,所提方法相比其他同类方法具有明显优势,刚性对齐后的平均每关节位置误差(PA‑MPJPE)最高降低了12.50个百分点,检测关节点百分比曲线下面积(AUC)最高提高了3.44个百分点。
Abstract:
In the task of 3D hand pose estimation from a single image in color, challenges such as occlusion and high self‑similarity of hand parts might lead to large prediction errors and unnatural hand structures. To address these issues, a graph convolution‑based 3D hand pose estimation method was firstly proposed. Visual features and 2D keypoint positions were extracted from the input image using Keypoint R‑CNN. These features were then fed into an improved adaptive kernel graph convolution module (AK_GraFormer). Subsequently, a residual‑connected AKGNN graph kernel was introduced to adaptively process graph‑structured data, thereby enhancing the model’s feature learning and representation. Finally, a dynamic training strategy was employed, which was monitored by a proposed evaluation metric, to optimize estimation performance. Experimental results on the HO‑3D_v3 and FreiHand datasets demonstrated that the proposed method outperformed existing approaches in monocular 3D hand pose estimation. Specifically, the procrustes‑aligned mean per joint position error (PA‑MPJPE) was reduced by up to 12.50 percentage points, and the area under the curve (AUC) of the percentage of correct keypoints metric was improved by up to 3.44 percentage points compared to state‑of‑the‑art methods.

参考文献/References:

[1] Sridhar S, Feit A M, Theobalt C, et al. Investigating the dexterity of multi‑finger input for mid‑air text entry[C]//Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems. New York: ACM, 2015: 3643‑3652.
[2] Oikonomidis I, Kyriazis N, Argyros A. Tracking the articulated motion of two strongly interacting hands[C]//Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2012: 1862‑1869.
[3] Tkach A, Pauly M, Tagliasacchi A. Sphere‑meshes for real‑time hand modeling and tracking[J]. ACM Transactions on Graphics, 2016, 35(6): 1‑11.
[4] Romero J, Tzionas D, Black M J. Embodied hands: modeling and capturing hands and bodies together[J]. ACM Transactions on Graphics, 2017, 36(6): 1‑17.
[5] Pavlakos G, Choutas V, Ghorbani N, et al. Expressive body capture: 3D hands, face, and body from a single image[C]//Proceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2019: 10967‑10977.
[6] Keskin C, Kıraç F, Kara Y E, et al. Hand pose estimation and hand shape classification using multi‑layered randomized decision forests[C]//Proceedings of the 12th European conference on Computer Vision. New York: ACM, 2012: 852‑863.
[7] Tompson J, Stein M, Lecun Y, et al. Real‑time continuous pose recovery of human hands using convolutional networks[J]. ACM Transactions on Graphics, 2014, 33(5): 1‑10.
[8] Pan Xiaoying, Li Shoukun, Wang Hao, et al. LGCANet: lightweight hand pose estimation network based on HRNet[J]. The Journal of Supercomputing, 2024, 80(13): 19351‑19373.
[9] Hoang D C, Xuan Tan P, Pham D L, et al. Efficient multimodal fusion for hand pose estimation with hourglass network[J]. IEEE Access, 2024, 12: 113810‑113825.
[10] Zhan Zhi, Luo Guang. Multiscale feature fusion network for monocular complex hand pose estimation[J]. Electronics Letters, 2023, 59(24): e13044.
[11] Panteleris P, Oikonomidis I, Argyros A. Using a single RGB frame for real time 3D hand pose estimation in the wild[C]//Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). Piscataway: IEEE, 2018: 436‑445.
[12] Doosti B, Naha S, Mirbagheri M, et al. HOPE‑Net: a graph‑based model for hand‑object pose estimation[C]//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 13618‑13627.
[13] Zhao Weixi, Wang Weiqiang, Tian Yunjie. GraFormer: graph‑oriented transformer for 3D pose estimation[C]//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 20406‑20415.
[14] Li Zhixin, Shang Fanqi, Huan Zhan, et al. Human activity recognition based on hybrid feature graph convolutional neural network[J]. Journal of Zhengzhou University (Engineering Science), 2024, 45(4): 46‑52.[李志新,尚繁琦,宦展,等. 基于混合特征图卷积神经网络的人体行为识别方法[J]. 郑州大学学报(工学版), 2024, 45(4):46‑52.]
[15] Cai Yujun, Ge Lihao, Liu Jun, et al. Exploiting spatial‑temporal relationships for 3D pose estimation via graph convolutional networks[C]//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2019: 2272‑2281.
[16] Aboukhaddra A T, Malik J, Robertini N, et al. Shape‑GraFormer: GraFormer‑based network for hand‑object reconstruction from a single depth map[J]. IEEE Access, 2024, 12: 124021‑124031.
[17] Zhuang Nan, Mu Yadong. Joint hand‑object pose estimation with differentiably‑learned physical contact point analysis[C]//Proceedings of the 2021 International Conference on Multimedia Retrieval. New York: ACM, 2021: 420‑428.
[18] Zhang Maomao, Li Ao, Liu Honglei, et al. Coarse‑to‑fine hand‑object pose estimation with interaction‑aware graph convolutional network[J]. Sensors, 2021, 21(23): 8092.
[19] Ma Shenglei, Li Jinghua, Kong Dehui, et al. 3D hand pose estimation based on double branches with multi‑scale attention[J]. Chinese Journal of Computers, 2023, 46(7): 1383‑1395.[马胜蕾,李敬华,孔德慧,等. 基于双分支多尺度注意力的手三维姿态估计[J]. 计算机学报, 2023, 46(7):1383‑1395.]
[20] Yang Wenji, Xie Liping, Qian Wenbin, et al. Coarse‑to‑fine cascaded 3D hand reconstruction based on SSGC and MHSA[J]. The Visual Computer, 2025, 41(1): 11‑24.
[21] He Kaiming, Gkioxari G, Dollár P, et al. Mask R‑CNN[C]//Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 2980‑2988.
[22] Ju Mingxuan, Hou Shifu, Fan Yujie, et al. Adaptive kernel graph neural network[C]//Proceedings of the Thirty‑Sixth AAAI Conference on Artificial Intelligence (AAAI‑22). Washington D C: AAAI, 2022:7051‑7058.
[23] Vasconcelos C, Birodkar V, Dumoulin V. Proper reuse of image classification features improves object detection[C]//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 2740‑2750.
[24] Hampali S, Sarkar S D, Lepetit V. HO‑3D_v3: improving the accuracy of hand‑object annotations of the HO‑3D dataset[PP/OL]. V1. arXiv (2021‑07‑02)[2025‑12‑19]. https://doi.org/10.48550/arXiv.2107.00887.
[25] Zimmermann C, Ceylan D, Yang Jimei, et al. Frei‑HAND: a dataset for markerless capture of hand pose and shape from single RGB images[C]//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2019: 813‑822.
[26] Yang Bing, Xu Chuyang, Yao Jinliang, et al. 3D hand pose estimation method based on monocular RGB images[J]. Journal of Zhejiang University (Engineering Science), 2025, 59(1): 18‑26.[杨冰,徐楚阳,姚金良,等. 基于单目RGB图像的三维手部姿态估计方法[J]. 浙江大学学报(工学版), 2025, 59(1):18‑26.]
[27] Chen Yujin, Tu Zhigang, Kang Di, et al. Model‑based 3D hand reconstruction via self‑supervised learning[C]//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2021: 10446‑10455.
[28] Yang Lixin, Li Kailin, Zhan Xinyu, et al. ArtiBoost: boosting articulated 3D hand‑object pose estimation via online exploration and synthesis[C]//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2022: 2740‑2750.
[29] Zhang Hongwen, Tian Yating, Zhang Yuxiang, et al. PyMAF‑X: towards well‑aligned full‑body model regression from monocular images[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(10): 12287‑12303.
[30] Duran E, Kocabas M, Choutas V, et al. HMP: hand motion priors for pose and shape estimation from video[C]//Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Piscataway: IEEE, 2024: 6341‑6351.
[31] Chen Ping, Chen Yujin, Yang Dong, et al. I2UV‑Hand‑Net: image‑to‑UV prediction network for accurate and high‑fidelity 3D hand mesh modeling[C]//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 12909‑12918.
[32] Lin K, Wang Lijuan, Liu Zicheng. Mesh graphormer[C]//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 12919‑12928.
[33] Liu Zuyan, Lin Gaojie, Wang Congyi, et al. HandMIM: pose‑aware self‑supervised learning for 3D hand mesh estimation[PP/OL]. V1. arXiv (2023‑07‑29)[2025‑12‑19]. https://doi.org/10.48550/arXiv.2307.16061.
[34] Pavlakos G, Shan Dandan, Radosavovic I, et al. Reconstructing hands in 3D with transformers[C]//Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2024: 9826‑9836.

相似文献/References:

[1]关昌珊,邴万龙,刘雅辉,等.基于图卷积网络的多特征融合谣言检测方法[J].郑州大学学报(工学版),2024,45(4):70.[doi:10.13705/ j.issn.1671-6833.2024.01.011]
 GUAN Changshan,BING Wanlong,LIU Yahui,et al.Multi-feature Fusion Rumor Detection Method Based on Graph Convolutional Network[J].Journal of Zhengzhou University (Engineering Science),2024,45(5):70.[doi:10.13705/ j.issn.1671-6833.2024.01.011]
[2]徐贞顺,张文豪,王振彪,等.融合多信息的图卷积实体对齐方法[J].郑州大学学报(工学版),2026,47(3):108.[doi:10.13705/j.issn.1671-6833.2026.03.010]
 XU Zhenshun,ZHANG Wenhao,WANG Zhenbiao,et al.Multiple Information Graph Convolutional Network Entity Alignment Method[J].Journal of Zhengzhou University (Engineering Science),2026,47(5):108.[doi:10.13705/j.issn.1671-6833.2026.03.010]
[3]张 震,刘 博,李 卓,等.一种面向交通流量预测的自适应时空图卷积网络[J].郑州大学学报(工学版),2026,47(5):68.[doi:10.13705/j.issn.1671-6833.2025.05.011]
 ZHANG Zhen,LIU Bo,LI Zhuo,et al.An Adaptive Spatial-Temporal Graph Convolutional Network for Traffic Flow Forecas[J].Journal of Zhengzhou University (Engineering Science),2026,47(5):68.[doi:10.13705/j.issn.1671-6833.2025.05.011]

更新日期/Last Update: 2026-09-03