- Research Article
- 10.1007/s00371-025-04235-7
Semantic-integrated multi-model fitting for real-time VSLAM in highly dynamic environments
- Dec 13, 2025
- The Visual Computer
- Tiantian Zhang + 3 more +3
Publications from 2021 to 2026
Showing 10 of 50 papers
Semantic-integrated multi-model fitting for real-time VSLAM in highly dynamic environments
Correlation and Development of A Maximum Bite Force Prediction Model Based on Handgrip Force and BMI in Young Healthy Adults.
To examine the associations among handgrip force (HF), maximum bite force (MBF) and body mass index (BMI) in individuals of the same age group across both genders. It also explores the laterality correlation between MBF and HF. Furthermore, to establish a simple approach for clinical MBF assessment, we aimed to develop a predictive model for MBF. In this cross-sectional study, 102 healthy young adults (51 males, 51 females) underwent MBF measurement via a dental bite force tester and HF via a digital dynamometer. BMI was calculated from height and weight. Spearman's correlation assessed variable relationships and laterality patterns; gender differences were analysed using Mann-Whitney U tests. A multiple regression model was constructed to predict MBF. MBF and HF were significantly correlated in both genders (p < 0.01). BMI showed a stronger influence on MBF and HF in females (p < 0.01). Laterality correlations between HF and MBF were also stronger in females. Males exhibited higher HF and BMI (p < 0.01), while MBF showed no significant gender difference (p = 0.536). HF and BMI are strong predictors for MBF. The developed predictive model offers a practical and objective tool for MBF assessment in healthy young adults. This model provides a simple, reliable method for clinical evaluation of MBF. ChiCTR2500097064.
Read moreEnhancing Glass Defect Detection with Diffusion Models: Addressing Imbalanced Datasets in Manufacturing Quality Control
Visual defect detection in industrial glass manufacturing remains a critical challenge due to the low frequency of defective products, leading to imbalanced datasets that limit the performance of deep learning models and computer vision systems. This paper presents a novel approach using Denoising Diffusion Probabilistic Models (DDPMs) to generate synthetic defective glass product images for data augmentation, effectively addressing class imbalance issues in manufacturing quality control and automated visual inspection. The methodology significantly enhances image classification performance of standard CNN architectures (ResNet50V2, EfficientNetB0, and MobileNetV2) in detecting anomalies by increasing the minority class representation. Experimental results demonstrate substantial improvements in key machine learning metrics, particularly in recall for defective samples across all tested deep neural network architectures while maintaining perfect precision. The most dramatic improvement was observed in ResNet50V2’s overall classification accuracy, which increased from 78% to 93% when trained with the augmented data. This work provides a scalable, costeffective approach to enhancing automated defect detection in glass manufacturing that can potentially be extended to other industrial quality assurance systems and industries with similar class imbalance challenges.
Read moreRAVE for Speech: Efficient Voice Conversion at High Sampling Rates
Voice conversion has gained increasing popularity within the field of audio manipulation and speech synthesis. Often, the main objective is to transfer the input identity to that of a target speaker without changing its linguistic content. While current work provides high-fidelity solutions they rarely focus on model simplicity, high-sampling rate environments or stream-ability. By incorporating speech representation learning into a generative timbre transfer model, traditionally created for musical purposes, we investigate the realm of voice conversion generated directly in the time domain at high sampling rates. More specifically, we guide the latent space of a baseline model towards linguistically relevant representations and condition it on external speaker information. Through objective and subjective assessments, we demonstrate that the proposed solution can attain levels of naturalness, quality, and intelligibility comparable to those of a state-of-the-art solution for seen speakers, while significantly decreasing inference time. However, despite the presence of target speaker characteristics in the converted output, the actual similarity to unseen speakers remains a challenge.
Read moreBoosting Human Pose Estimation via Heatmap Refinement
Human pose estimation based on heatmap regression has achieved significant success in recent years. However, the semantic ambiguity caused by traditional hand-crafted heatmaps seriously affects the model performance. Specifically, hand-crafted heatmaps generated with a fixed Gaussian kernel are semantically misaligned. Various Gaussian covered areas for keypoints with the same type may cause model learning confusion. In this paper, we focus on learnable heatmap generation and propose a refined heatmap generator (RHG) to boost human pose estimation. First, we propose a joint training framework to connect the human pose estimator and RHG for end-to-end training. It employs a joint loss function to learn intermediate representations of the network and dataset. Second, RHG takes annotated dotpoints as input and utilizes scale-aware heatmaps as regression targets to deal with the scale variation. Scale-aware heatmaps are generated by adjusting Gaussian covered areas with geometric priors. Experimental results show that our method achieves 72.0%AP on COCO test-dev2017 and 74.0%AP on CrowdPose dataset, respectively, outperforming state-of-the-art methods.
Read moreSADNet: Generating Immersive Virtual Reality Avatars by Real-time Monocular Pose Estimation
Generating immersive virtual reality avatars is a challenging task in VR/AR applications, which maps physical human body poses to avatars in virtual scenes for an immersive user experience. However, most existing work is time-consuming and limited by datasets, which does not satisfy immersive and real-time requirements of VR systems. In this paper, we aim to generate 3D real-time virtual reality avatars based on a monocular camera to solve these problems. Specifically, we first design a self-attention distillation network (SADNet) for effective human pose estimation, which is guided by a pre-trained teacher. Secondly, we propose a lightweight pose mapping method for human avatars that utilizes the camera model to map 2D poses to 3D avatar keypoints, generating real-time human avatars with pose consistency. Finally, we integrate our framework into a VR system, displaying generated 3D pose-driven avatars on Helmet-Mounted Display devices for an immersive user experience. We evaluate SADNet on two publicly available datasets. Experimental results show that SADNet achieves a state-of-the-art trade-off between speed and accuracy. In addition, we conducted a user experience study on the performance and immersion of virtual reality avatars. Results show that pose-driven 3D human avatars generated by our method are smooth and attractive.
Read moreNavigating Financial Transactions in the Metaverse: Risk Analysis, Anomaly Detection, and Regulatory Implications
Blockchain technology has emerged as a disruptive force in the realm of finance, offering decentralized and transparent mechanisms for conducting financial transactions. This paper explores the landscape of blockchain-based financial transactions, focusing on risk analysis, anomaly detection, regulatory frameworks, and ethical considerations. Drawing on interdisciplinary insights from finance, computer science, economics, law, and ethics, the study investigates the opportunities and challenges presented by blockchain finance. Leveraging quantitative analysis, machine learning algorithms, case studies, and regulatory reviews, the research sheds light on the complexities of blockchain ecosystems. Key findings include the importance of robust risk management strategies, the role of anomaly detection in safeguarding financial integrity, and the evolving regulatory landscape surrounding blockchain transactions. The study identifies gaps in current research and proposes avenues for future investigation, emphasizing the need for interdisciplinary approaches to address the multifaceted challenges of blockchain-based finance. Ultimately, this research aims to inform stakeholders about the implications of blockchain technology in financial transactions and foster responsible innovation and sustainable development in digital finance ecosystems.
Read moreThe Journey from Non-Immersive to Immersive Multi-user Applications in Mental Health Care: Systematic Review (Preprint)
BACKGROUND Over the past 25 years, the development of multi-user applications has seen significant advancements and challenges. The technological development in this field has emerged from simple chatrooms, through videoconferencing tools to the creation of complex, interactive, and often multisensory virtual worlds. These multi-user technologies have gradually found their way into mental health care, where they are used in both dyadic counseling and group interventions. However, some limitations in hardware capabilities, user experience designs, and scalability may have hindered the effectiveness of these applications. OBJECTIVE The present systematic review aimed at summarizing the progress made and the potential future directions in this field while evaluating various factors and perspectives relevant to remote multi-user interventions. METHODS The systematic review was performed based on Web of Science (WoS) and PubMed database search covering articles in the English language published from January 1999 to March 2024 related to multi-user mental health interventions. Several inclusion and exclusion criteria were determined before and during the records screening process performed in several steps. RESULTS We have identified 49 records exploring the multi-user applications in mental health care, ranging from text-based interventions to interventions set in fully immersive environments. The number of publications exploring this topic is growing since 2015, with a large increase during COVID-19 pandemic. The majority of digital interventions were delivered in a form of video-conferencing, with only a few implementing immersive environments. The studies utilized professional or peer supported group interventions or a combination of both approaches. The research studies targeted diverse groups and topics, from nursing mothers to psychiatric disorders or various minority groups. Most group sessions happened weekly, or in case of the peer-suport groups, often with flexible schedule. CONCLUSIONS We have identified many benefits to multi-user digital interventions for mental healthcare. These approaches provide distributed, always available and affordable peer support that can be used to deliver necessary help to people living outside of areas where in-person interventions are easily available. While immersive virtual environments have become a common tool in many areas of psychiatric care, such as exposure therapy, our results suggest that this technology in multi-user settings is still in its early stages. Most identified studies investigated mainstream technologies, such as video conferencing or text-based support, substituting immersive experience for convenience and ease of use. While many studies discuss useful features of virtual environments in group interventions, such as anonymity or stronger engagement with the group, we discuss persisting issues with these technologies, which currently prevent their full adoption. CLINICALTRIAL N/A
Read moreWhat Is Medical Extended Reality? A Taxonomy Defining the Current Breadth and Depth of an Evolving Field.
Medical extended reality (MXR) has emerged as a dynamic field at the intersection of health care and immersive technology, encompassing virtual, augmented, and mixed reality applications across a wide range of medical disciplines. Despite its rapid growth and recognition by regulatory bodies, the field lacks a standardized taxonomy to categorize its diverse research and applications. This American Medical Extended Reality Association guideline, authored by the editorial board of the Journal of Medical Extended Reality, introduces a comprehensive taxonomy for MXR, developed through a multidisciplinary and international collaboration of experts. The guideline seeks to standardize terminology, categorize existing work, and provide a structured framework for future research and development in MXR. An international and multidisciplinary panel of experts was convened, selected based on publication track record, contributions to MXR, and other objective measures. Through an iterative process, the panel identified primary and secondary topics in MXR. These topics were refined over several rounds of review, leading to the final taxonomy. The taxonomy comprises 13 primary topics that jointly expand into 180 secondary topics, demonstrating the field's breadth and depth. At the core of the taxonomy are five overarching domains: (1) technological integration and innovation; (2) design, development, and deployment; (3) clinical and therapeutic applications; (4) education, training, and communication; and (5) ethical, regulatory, and socioeconomic considerations. The developed taxonomy offers a framework for categorizing the diverse research and applications within MXR. It may serve as a foundational tool for researchers, clinicians, funders, academic publishers, and regulators, facilitating clearer communication and categorization in this rapidly evolving field. As MXR continues to grow, this taxonomy will be instrumental in guiding its development and ensuring a cohesive understanding of its multifaceted nature.
Read moreHolokausta komemorācijas un izglītības pārnese digitālajā vidē: atmiņas digitalizācija, datorspēles un virtuālā realitāte
Komunikcija par sareto kultrvsturisko mantojumu nenorit vien fiziskaj pasaul. Paralli muzeju ekspozcijm un arhviem virtul pasaule piedv arvien jaunus veidus, k runt, spriest un atcerties pagtnes notikumus
Read more