- Research Article
381
- 10.1016/j.cviu.2021.103225
Deep 3D human pose estimation: A review
- May 24, 2021
- Computer Vision and Image Understanding
- Jinbao Wang + 6 more +6
Deep 3D human pose estimation: A review
Museums are key institutions for cultural communication and public education, and their operating concept is shifting from exhibit-centered to experience-centered. As expectations for exhibition experience rise, museum fatigue has become a major constraint on visitors. Existing studies rely on questionnaires and other subjective measures, which makes it difficult to locate fatigue in specific spaces. At the same time, body pose detection and fatigue recognition techniques remain hard to apply in museums because of complex spatial configurations and dense visitor flows. Effective methods for quantifying and mitigating museum fatigue are still lacking. This study proposes a contact-free sensing scheme based on computer vision and builds a coupled analytical framework with three stages: Human Pose Estimation (HPE) for visitor posture detection, fatigue assessment, and fatigue mitigation. A Fatigue Index (FI) quantifies bodily fatigue. Applying this index to the exhibition space in both the baseline and adjusted configurations guides the formulation of mitigation strategies and shows a consistent reduction in FI, which indicates that the adopted measures are effective. The proposed approach establishes a complete frame from fatigue quantification to fatigue mitigation, supports evaluation of exhibition space design, and provides theoretical and methodological support for future improvements to museum experience.
Deep 3D human pose estimation: A review
Deep 3D human pose estimation: A review
Quantitative assessment of upper limb muscle fatigue depending on the conditions of repetitive task load
Quantitative assessment of upper limb muscle fatigue depending on the conditions of repetitive task load
Deep Full-Body HPE for Activity Recognition from RGB Frames Only
Human Pose Estimation (HPE) is defined as the problem of human joints’ localization (also known as keypoints: elbows, wrists, etc.) in images or videos. It is also defined as the search for a specific pose in space of all articulated joints. HPE has recently received significant attention from the scientific community. The main reason behind this trend is that pose estimation is considered as a key step for many computer vision tasks. Although many approaches have reported promising results, this domain remains largely unsolved due to several challenges such as occlusions, small and barely visible joints, and variations in clothing and lighting. In the last few years, the power of deep neural networks has been demonstrated in a wide variety of computer vision problems and especially the HPE task. In this context, we present in this paper a Deep Full-Body-HPE (DFB-HPE) approach from RGB images only. Based on ConvNets, fifteen human joint positions are predicted and can be further exploited for a large range of applications such as gesture recognition, sports performance analysis, or human-robot interaction. To evaluate the proposed deep pose estimation model, we apply it to recognize the daily activities of a person in an unconstrained environment. Therefore, the extracted features, represented by deep estimated poses, are fed to an SVM classifier. To validate the proposed architecture, our approach is tested on two publicly available benchmarks for pose estimation and activity recognition, namely the J-HMDBand CAD-60datasets. The obtained results demonstrate the efficiency of the proposed method based on ConvNets and SVM and prove how deep pose estimation can improve the recognition accuracy. By means of comparison with state-of-the-art methods, we achieve the best HPE performance, as well as the best activity recognition precision on the CAD-60 dataset.
Read moreAction Recognition based on Human Pose Estimation
In computer vision, human pose estimation and the following action recognition are active research topics, which have a wide application in human computer interaction, smart homes, athlete assistance training, etc. Driven by these applications, numerous algorithms and models are proposed in the recent years, among which deep learning becomes dominant. In this study, we summarize the recent progress of deep learning in human pose estimation and action recognition. We also point out some research challenges and future research directions.
Read moreDGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation
Monocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a two-stage pipeline by establishing correspondences and utilizing a RANSAC algorithm to calculate 6 degrees-of-freedom (6DoF) pose. Recent works try to integrate differentiable RANSAC algorithms to achieve an end-to-end 6D pose estimation. However, most of them hardly consider the geometric features in 3D space, and ignore the topology cues when performing differentiable RANSAC algorithms. To this end, we proposed a Depth-Guided Edge Convolutional Network (DGECN) for 6D pose estimation task. We have made efforts from the following three aspects: 1) We take advantages of estimated depth information to guide both the correspondences-extraction process and the cascaded differentiable RANSAC algorithm with geometric information. 2) We leverage the uncertainty of the estimated depth map to improve accuracy and robustness of the output 6D pose. 3) We propose a differentiable Perspective-n-Point(PnP) algorithm via edge convolution to explore the topology relations between 2D-3D correspondences. Experiments demonstrate that our proposed network outperforms current works on both effectiveness and efficiency.
Read moreMonocular human pose estimation: A survey of deep learning-based methods
Monocular human pose estimation: A survey of deep learning-based methods
Using deep learning to detect upper limb compensation in individuals post-stroke using consumer-grade webcams—A feasibility study
As societies age, the number of individuals experiencing stroke increases, necessitating more effective rehabilitation strategies. Over half of stroke survivors suffer from upper limb impairments, making assessments of sensory-motor function crucial for both improving interventions and tracking progress. Ideally, such assessments could also be performed at home without requiring a therapist's presence. Advances in computer vision and human pose estimation allow for human movement analysis using consumer-grade cameras. This study investigates whether a single webcam, combined with human pose estimation and deep learning algorithms, can automatically detect compensatory movements in persons with stroke performing a drinking task. Twenty participants with stroke with mild to moderate upper limb impairment were recruited. Each participant performed multiple repetitions of the drinking task while being recorded by multiple cameras and an optical motion capture system (OMC) for kinematic ground truth. The videos were labeled by therapists to indicate the presence or absence of compensatory movements. Human poses were extracted from the videos using MediaPipe, and deep learning models were trained to predict these compensatory movements based on MediaPipe keypoints. Several factors affecting compensation detection accuracy were evaluated. Models trained on raw MediaPipe keypoints for inter-person compensation detection failed to generalize, achieving accuracy around 50%. Using custom features instead of raw keypoints improved the accuracy to 70%. In contrast, intraperson classification achieved high accuracy, typically exceeding 90%. Using OMC data significantly improved classification accuracy compared to using MediaPipe keypoints. Camera angle had an effect on accuracy, and convolutional neural networks outperformed long short-term memory networks. Generalizing models remain limited by (1) the measurement uncertainty of human pose estimation and (2) insufficient data representing the full spectrum of compensatory strategies (3) accurate compensation labels. The results demonstrate that deep learning approaches can differentiate between compensatory and non-compensatory movements when movement representations are sufficiently accurate. Future work should improve pose estimation and expand labeled datasets to better reflect the stroke population. While general models are limited in accuracy, personalized models using consumer cameras can support home-based rehabilitation. This digitalized assessment approach has the potential to quantify recovery progress throughout the continuum of care.
Read moreHuman Pose Estimation Using OpenCV
Human pose estimation is a crucial task in computer vision, involving the detection and tracking of key body joints in images or videos. This technology has numerous applications, including activity recognition, augmented reality, and human-computer interaction.This paper presents an approach to human pose estimation using OpenCV in combination with deep learning-based models such as OpenPose or MediaPipe Pose. The method leverages a pre- trained neural network that detects key body landmarks, such as the head, shoulders, elbows, and knees, from RGB images. OpenCV provides efficient image processing tools to preprocess input frames, detect poses, and visualize the estimated skeletal structure.Our implementation focuses on real time processing by optimizing inference speed while maintaining high accuracy. The proposed system is tested on various datasets to evaluate its robustness under different lighting conditions and human postures. Results demonstrate the effectiveness of OpenCV-based pose estimation in achieving reliable skeletal tracking with minimal computational overhead. This study highlights the potential of OpenCV in real-time human pose estimation and its applications in fitness tracking, gesture recognition, and motion analysis. Future work includes improving model accuracy, reducing latency, and integrating pose estimation with advanced AI- driven applications.
Read moreReliable feature point detection and object pose estimation using photometric quasi-invariant SIFT
Object pose estimation from stereo images with unknown correspondence is a thoroughly studied problem in the computer vision and robot engineering literatures. Especially, it is important to detect the desirable corresponding points from images for object pose estimation. For this, many approaches have been proposed. Among them, the local feature descriptor, which describe the feature points that are robust to image deformations in an object or image, is one of the most promising approaches that has been applied to the stable feature detection problem successfully. Although any descriptors including the SIFT represent superior performance, these are based on luminance information rather than color information thereby resulting in instability to photometric variations such as shadows, highlights, and illumination changes. Therefore, we propose a novel method which extracts the interest points that are insensitive to both geometric and photometric variations in order to estimate more accurate and desirable object pose. In this method, we use photometric quasi-invariant features based on the dichromatic reflection model in order to achieve photometric invariance, and the SIFT is used for geometric invariance as well. The performance of the proposed method is evaluated with other local descriptors. Experimental results show that our method gives similar performance or outperforms them with respect to various imaging conditions. Finally, we estimate object pose by using the features extracted via the proposed method.
Read moreCanonical Shape Reconstruction With SE(3) Equivariance Learning for Weakly-Supervised Object Pose Estimation
6D object pose estimation from a single RGB-D image is a fundamental problem in computer vision and robot manipulation. Despite recent advancements, existing methods still suffer several limitations. First of all, the object shape representation extracted from the depth map is often less expressive because the object point cloud parsed from the depth map is highly incomplete due to the object self-occlusion and noisy due to the sensor artifacts. This shape representation issue further intensifies when lacking sufficient labeled data for model training, which unfortunately is another typical problem for object pose estimation considering the heavy annotation cost for real-world pose labeling. In this study, we propose to tackle the above issues in a unified way. First, we enhance the object shape representation from the partial point cloud with a novel canonical shape reconstruction module, in which an implicit canonical frame is established by incorporating the SE(3) equivariance, achieving implicit feature alignment of the partial point cloud inputs, leading to robust shape recovery. Second, based on the enhanced object representation, we further utilize the de-canonicalized and pose-dependent completed object shape as the training signal, and develop a novel weakly-supervised learning framework to leverage both labeled synthetic data and unlabeled real data to train the pose estimation model in a label-efficient way. Extensive experiments on three widely used benchmarks demonstrate the effectiveness, and superiority of our framework over state-of-the-art methods.
Read moreCross-Modal Vision-Language Model for High-Precision Human Pose Estimation
Human pose estimation (HPE) is a critical task in computer vision, with applications in behavior monitoring, human-computer interaction, action recognition, and online learning. However, practical HPE applications often face challenges such as object occlusion, crossed limbs, blurred backgrounds, and complex backgrounds, which significantly impact the accuracy of keypoint localization. Traditional handcraft feature-based approaches and convolutional neural network (CNN)-based methods have limitations in handling these challenges. This paper proposes a novel Cross-Modal High-Precision Human Pose Estimation (CMP-HPE) method that leverages a cross-modal transformer to integrate visual and language information, enhancing the robustness and accuracy of HPE. The proposed model comprises four main modules: visual tokens construction, language tokens construction, cross-modal transformer, and marker-based prediction. Experiments on the MSCOCO and MPII datasets demonstrate the superior performance of our method, achieving significant improvements over state-of-the-art techniques. The source Python code will be available upon request.
Read moreA Monadic and Effective Frame Work for Single Human Pose Estimation of 2D Images and Videos
Human pose estimation (HPE) is most dominant research areas in Computer Vision (CV) field. This technology will have huge implications. The significant applications include human computer interaction, activity recognition, video surveillance, motion recognition, etc. Different types of pose estimation are used for estimating the number of persons who are tracked. In this paper we are focusing on single 2D pose estimation. Present trends of pose estimation uses CNN based architectures for HPE and post statistical methods. In this work we propose a monadic frame work for both HPE and post processing into a single stage. It produces a human pose skeleton for a single two-dimensional image and videos. Here, take different modalities such as static image, static video and live video to generate different skeletons for SHPE. In this work we use pre-trained tensor flow based CNN model for pose estimation and also it produce better results compared to state of art in terms of performance metrics as PCP@.5, PCK@.5, and PCKh@.2.
Read moreInteraction between the leg recovery test and subjective measures of fatigue in handball players: short-, mid-, and long-term assessment
BackgroundThe physical and mental demands of handball during training or competition often lead to fatigue which can impair performance. Many attempts have been made to assess the level of fatigue in athletes either by objective (neuromuscular performance) or subjective (questionnaires) measures, however, their interplay over short-, mid-, and long-term periods is currently unknown. Knowledge about both types of assessments is important as load management by coaches is traditionally based on direct adjustments following a training session, adjustments of content structure of training weeks between games, as well as adjustments of load management over the entire competitive season. Thus, this study aimed to investigate the interplay between objective and subjective fatigue measures at multiple test times throughout a handball season.MethodsA total of 100 highly trained (Tier level 3) adolescent or young adult team handball players (23 females) took part in the study. The parameters tested were the Leg Recovery Test (LRT score) which is based on the countermovement jump height (CMJ) and was assessed by a commercial wristwatch (Polar Vantage V2) as an objective measure of neuromuscular fatigue. Additionally, on a subjective level, questionnaire-based athlete self-report measures, specifically the Perceived Recovery Status Scale (PRSS) and the Short Scale of Recovery and Strain (KEB) were assessed. We used non-parametric tests to detect differences between relevant test time points (short-term: immediately following one handball-specific training session, i.e., from T0 to T1; mid-term: over the course of three consecutive training days, i.e., from T0 to T2; long-term: over the course of 8 months of training, i.e., from T0 to T12) and linear mixed models to evaluate the interplay between objective (LRT score) and subjective (KEB score and PRSS score) measures of fatigue across one season.ResultsNon-parametric tests showed that CMJ height (p = .012) and the KEB (p < .001) were higher at T1 compared to T0 for the short-term assessment. Over the course of three consecutive training days (i.e., mid-term assessment), the CMJ height score decreased (T0 to T2: p < .001; T1 to T2: p = .018) and the PRSS score (T0 to T2: p < .001; T1 to T2: p = .003) increased. Linear mixed models revealed no significant effects of KEB or PRSS score on LRT score (i.e., CMJ height) for the short- and mid-term assessments. In terms of the long-term assessments, we detected no general direct or interaction effects of PRSS score, workload, and test time point on LRT score, except for an interaction between PRSS score and workload on LRT score (p = .032), which indicates a workload-dependent association between PRSS and the objective fatigue measure (LRT score).ConclusionAthlete self-reported measures of fatigue indicated significantly higher cumulative fatigue after both short- and mid-term periods, whereas this increase was observed in the LRT score only during the mid-term period. Furthermore, the absence of a relationship between the objective and subjective measures of fatigue during short- and mid-term periods suggests that these measures assess distinct types of fatigue. In the long-term assessments, the significant interaction between the PRSS score and workload on the LRT score suggests that higher workloads are associated with an increased correlation between subjective (PRSS score) and objective (LRT score) measures of fatigue. This indicates that perceived fatigue may be a more sensitive indicator of fatigue, which can be managed to maintain high levels of neuromuscular performance (LRT score). However, with higher workloads (>10 h per week), associations between the objective and subjective measures become apparent, suggesting that workload serves as a common factor influencing overall fatigue.
Read morePosturographic Balance's Validity in Mental and Physical Fatigue Assessment Among Cadet Pilots.
BACKGROUND: Postural control is adversely affected by mental and physical fatigue, but its validity in fatigue assessment has not been investigated systemically among pilots. We explored the correlations of posturographic balance with physiological and psychological signals among cadet pilots.METHODS: In experiment 1, 37 cadet pilots performed a posturographic balance test, heart rate variability (HRV), and profile of mood states (POMS) during 40 h of sleep deprivation. For experiment 2, physiological signals of 60 subjects, including breathing rate (BR), systolic blood pressure (SBP), and heart rate (HR) were measured under the effects of physical fatigue. Then correlations with a mental and physical fatigue index based on effective posturographic parameters with those subjective and objective methods were analyzed by linear regression.RESULTS: The mental fatigue index correlated linearly with the depression score of the POMS (r = 0.212), standard deviation of normal to normal beats (r = 0.286), and square root of the mean differences of successive beat intervals (r = 0.207). Meanwhile, linear correlations with frequency-domain parameters of HRV such as total power, low frequency power, and high frequency power were also statistically significant. With the increase in the physical fatigue index, physiological signals such as SBP (r = 0.300), HR (r = 0.349), and BR (r = 0.266) increased linearly.CONCLUSIONS: Impairment of postural stability can reflect the aggravation of mental and physical fatigue among cadet pilots, which provides a potential method for assessing fatigue level before flight tasks and preventing errors by pilots.Cheng S, Sun J, Ma J, Dang W, Tang M, Hui D, Zhang L, Hu W. Posturographic balance's validity in mental and physical fatigue assessment among cadet pilots. Aerosp Med Hum Perform. 2018; 89(11):961-966.
Read moreAnd yet it works!
The purpose of this book is to help the reader understand the significance of public education in alcohol and drug prevention and to encourage those involved in alcohol and drug prevention to leverage communications means and to develop their media savvy. Our main point is that public education does work and that research criticising public education is wrong. We also seek to encourage exploration of new methods for evaluating and studying public education so as to consider how the importance and impact of communications in public alcohol and drug education could be enhanced.
Read more