Font Size: a A A

Research On 3D Human Pose Estimation Based On Monocular Camera

Posted on:2024-02-08Degree:MasterType:Thesis
Country:ChinaCandidate:X ChenFull Text:PDF
GTID:2568307106467644Subject:Information and Communication Engineering
Abstract/Summary:
3D human pose estimation based on monocular camera has far-reaching significance in reducing the use cost and expanding the range of use scenes,which is the mainstream research hotspot at present.At present,the mainstream monocular 3D human pose estimation uses a two-stage method.In the first stage,2D human pose estimation is used to obtain 2D human node coordinates,but the accuracy is limited.In the second stage,the 3D human pose estimation network was used to return the 3D human node coordinates,but the depth information was still ambiguous.Aiming at the problem of the first stage,a trajectory correction algorithm based on Kalman filter is proposed to make up for the accuracy problem.Aiming at the problem of the second stage,a 2D-3D network based on the human topological motion tree structure is proposed to improve the depth information and reduce the error.Through the proposed two-stage improvement method,a more accurate monocular 3D human pose estimation is achieved,which has important research significance.The main work and innovation of this paper are as follows:(1)2D human pose tracking correction algorithm based on Kalman filter is designed.In order to overcome the loss of key points in the occlusion problem caused by 2D body estimation,a set of tracking algorithm is adopted.If the current point coordinates are the same as the Kalman predicted value,the result of 2D body pose estimation is used.If the current coordinates are lost or the gap between them and the current Kalman predicted value is too large,the predicted point of Kalman filter is used to replace the original detection result for tracking.Through this algorithm design,the detection results of 2D pose estimation can be improved to a large extent.At the same time,as the input of the second stage,more reliable 2D key point coordinates of human body are provided.(2)Based on the human motion tree structure,a feature fusion network based on position enhancement is proposed.In order to resist the interference of global motion(such as camera translation)and solve the problem of declining network learning ability caused by different distribution of the input and output terminals of the network,the position enhancement algorithm is adopted for the position information of the input terminal.The experimental results show that the error is reduced by 0.7mm compared with the reference line.Based on the time convolution structure,a time-enhanced feature fusion network is proposed.In order to solve the performance degradation of local motion changes,a time enhancement algorithm was proposed to learn the influence of other postures on the current posture and other postures description-driven network,and to return to more accurate 3D human posture coordinates.The experimental results show that the error is reduced by 1mm compared with the reference line.(3)Based on the integrated optimization strategy of network,the encoder,decoder and feature fusion module of the feature fusion network are respectively optimized in three stages by making full use of the information exchange within the human body group and the information dependence between the groups on the basis of the spatial enhancement and time enhancement algorithms.In the first two stages,coding and fusion were independently optimized to enable each group to independently extract the time and space information related to pose and share the connection between groups.Finally,network integration was completed through fine tuning in the third stage.Experimental results showed that the error was reduced by 4mm compared with the reference line.
Keywords/Search Tags:monocular 3D human pose estimation, Kalman filter, spatial enhancement, time enhancement
Related items