- Research Article
15
- 10.11834/jig.230020
Human-computer interaction for virtual-real fusion
- Jan 01, 2023
- Journal of Image and Graphics
- Jianhua Tao + 5 more +5
Human-computer interaction for virtual-real fusion
The European research project INT-MANUS embedded in the I*PROMS European network of excellence addresses the increasing demand for flexibility and adaptivity, which is summarised by rapid reconfiguration of complete factories, flexible reaction to new demands as well as related aspects in human computer interaction (HCI), software, and production systems. The project's main goal has been to develop a new technology for the production plants of the future: the Smart Connected Control Platform (SCCP). This platform allows controlling a factory with the help of an open distributed and learning agent platform that integrates machines, robots, and human personnel. It offers an enterprise service bus like concept for dynamic and decentrally controlled production systems, which flexibly connects machines and IT systems like robotic transport systems, terminals, mobile control systems, etc.
Human-computer interaction for virtual-real fusion
Human-computer interaction for virtual-real fusion
Object Acquisition and Selection in Human Computer Interaction Systems: A Review
Object acquisition and selection are two important functions performed in most of the human computer interaction (HCI) systems. Various techniques are devised by the researchers to perform these operations and the selection of a combination of object acquisition and selection techniques for a particular HCI system has become a research issue, especially, when the user of these systems are differently abled persons. This paper presents a review on object acquisition and selection techniques used in HCI systems. It starts with the introduction to HCI systems, gives an overview on object acquisition & selection techniques, feedback modes, performing mouse analogous actions using eye blinks, the applications of HCI systems, and finally discusses challenges and issues related to these techniques.
Read moreIntelligence methods of multi-modal information fusion in human-computer interaction
We first introduce the concepts of single-modal information processing and multi-modal information fusion in cognitive science. Some classical multi-modal information fusion models and their computer implementations in history are also explained. Under the conditions that each channels information can be obtained, and their features could be unified representation synchronously, the fusion of multi-modal information can be transformed into classification or regression problems. For practical human-computer interaction systems, the performance of multi-modal information fusion largely relies on the accuracy of the single-modal information identification and the design of the interactive system. We present a practical example of multi-modal information fusion system, and discuss its performances on human computer interaction. Finally, the possible and important development trends for multi-modal human-computer interaction techniques and systems are discussed.
Read moreBuilding it better: Applying human–computer interaction and persuasive system design principles to a monetary limit tool improves responsible gambling
Building it better: Applying human–computer interaction and persuasive system design principles to a monetary limit tool improves responsible gambling
Read moreFabrication of Laser-Induced Graphene Based Flexible Sensors Using 355 nm Ultraviolet Laser and Their Application in Human-Computer Interaction System.
In recent years, flexible sensors based on laser-induced graphene (LIG) have played an important role in areas such as smart healthcare, smart skin, and wearable devices. This paper presents the fabrication of flexible sensors based on LIG technology and their applications in human-computer interaction (HCI) systems. Firstly, LIG with a sheet resistance as low as 4.5 Ω per square was generated through direct laser interaction with commercial polyimide (PI) film. The flexible sensors were then fabricated through a one-step method using the as-prepared LIG. The applications of the flexible sensors were demonstrated by an HCI system, which was fabricated through the integration of the flexible sensors and a flexible glove. The as-prepared HCI system could detect the bending motions of different fingers and translate them into the movements of the mouse on the computer screen. At the end of the paper, a demonstration of the HCI system is presented in which words were typed on a computer screen through the bending motion of the fingers. The newly designed LIG-based flexible HCI system can be used by persons with limited mobility to control a virtual keyboard or mouse pointer, thus enhancing their accessibility and independence in the digital realm.
Read moreTalking With Your Hands: Scaling Hand Gestures and Recognition With CNNs
The use of hand gestures provides a natural alternative to cumbersome interface devices for Human-Computer Interaction (HCI) systems. As the technology advances and communication between humans and machines becomes more complex, HCI systems should also be scaled accordingly in order to accommodate the introduced complexities. In this paper, we propose a methodology to scale hand gestures by forming them with predefined gesture-phonemes, and a convolutional neural network (CNN) based framework to recognize hand gestures by learning only their constituents of gesture-phonemes. The total number of possible hand gestures can be increased exponentially by increasing the number of used gesture-phonemes. For this objective, we introduce a new benchmark dataset named Scaled Hand Gestures Dataset (SHGD) with only gesture-phonemes in its training set and 3-tuples gestures in the test set. In our experimental analysis, we achieve to recognize hand gestures containing one and three gesture-phonemes with an accuracy of 98.47% (in 15 classes) and 94.69% (in 810 classes), respectively. Our dataset, code and pretrained models are publicly available.
Read moreProduct Design for the Elderly Based on Human-Computer Interaction in the Era of Big Data
Since 21st century, social aging trend has become more and more serious. It needs accurately identify the nursing needs of elderly disabled people, who are characterized by inconvenient movement and unclear speech. How to solve these problems has become a key point in the field of elderly care and medical care. In response to this problem, this research has designed a human-computer interactive gesture recognition system for elderly nursing beds in the context of big data. The recognition rate of the fusion feature + support vector machine (SVM) classifier adopted in this study is higher than 90% for each category of gesture. On the test set, this method has an average recognition rate of 96.35%, which is much higher than that of single feature + SVM classifier. While other methods’ recognition rate is lower than 90%. The recognition rate of tag c (nursing bed posture turning left) with obvious gesture feature information is as high as 99.28%, and that of tag h (nursing bed posture bedpan lowering) with weak gesture feature is 93.65%. The human-computer interaction system has well realized the recognition intention of user’s dynamic and static gestures, achieved the goal set by the research, and the interaction form is natural and reliable. In the later research, we can further realize a more comprehensive, accurate and natural human-computer interaction product design through the multi-channel joint decision-making scheme to meet the needs of the elderly.
Read moreReal Time Driver Drowsiness Detection Based on Driver’s Face Image Behavior Using a System of Human Computer Interaction Implemented in a Smartphone
The main reason for motor vehicular accidents is the driver drowsiness. This work shows a surveillance system developed to detect and alert the vehicle driver about the presence of drowsiness. It is used a smartphone like small computer with a mobile application using Android operating system to implement the Human Computer Interaction System. For the detection of drowsiness, the most relevant visual indicators that reflect the driver’s condition are the behavior of the eyes, the lateral and frontal assent of the head and the yawn. The system works adequately under natural lighting conditions and no matter the use of driver accessories like glasses, hearing aids or a cap. Due to a large number of traffic accidents when driver has fallen asleep this proposal was developed in order to prevent them by providing a non-invasive system, easy to use and without the necessity of purchasing specialized devices. The method gets 93.37% of drowsiness detections.
Read moreAn efficient interpretation of hand gestures to control smart interactive television
In the recent era of smart world, smarter technologies are gaining focus on human computer interaction (HCI) systems, and traditional ways like remote control, mouse, keyboard, etc. are becoming less popular. This paper presents a framework of simple yet efficient approach for future applications of the new age smart and intelligent technologies that shall enhance the HCI specifically for smart television. Hand gesture recognition (HGR)-based model is proposed for wireless control of smart interactive television (SITV), which includes controlling volume and selecting channels. The proposed framework undergoes three steps: 1) the hand gesture of the person is detected by using shape, colour and skin similarity; 2) the extracted features are classified by using rule-based classification and a gesture code is generated; 3) the classified gesture is interpreted by a novel interpreter. The performance of the proposed framework is evaluated with different people's hand gestures and compared with the techniques of others.
Read moreModel Adaptation Approach to Speech Synthesis with Diverse Voices and Styles
In human computer interaction and dialogue systems, it is often desirable for text-to-speech synthesis to be able to generate natural sounding speech with an arbitrary speaker's voice and with varying speaking styles and/or emotional expressions. We have developed an average-voice-based speech synthesis method using statistical average voice models and model adaptation techniques for this purpose. In this paper, we describe an overview of the speech synthesis system and show the current performance with several experimental results.
Read moreHuman Resources and Knowledge Management Based on E-Democracy
Correspondence between everyday and scientific life is a blessing (Warren & Jahoda, 1966). Although virtual communities are nowadays widely expanded, research online is not fully developed yet, as methodological approaches are not designed specifically for online research. In addition, the results from the evaluation and the reports do not find an immediate space of use. As such, the researchers use methodologies that deal with online situations borrowing methods and techniques from the “real” ones. Although the adaptations have the same principles, there are limitations due to the virtual nature of the research. In addition, multi-disciplinary approaches characterize virtual communities as different fields interact, such as learning approaches, psychology of the individual and the masses, sociology, linguistics, communication studies, management, human computer interaction and information systems. As a result, there is no methodology that, solely used, could bring results for adequate evaluation and implementation of the results in the community. Due to this complexity, we suggest Real Time Research Methodology based on Time-Series Design to study process-based activities; Focus Groups Methodology and Forum Messages Discourse Analysis as two of the most vital parts in the use of a multi-method. The other parts will depend on the nature and culture of the selected virtual community. Both focus groups (FG) and Forum Messages Discourse Analysis are referred as Extraction Group Research Methodology, or X-Groups. The reason for using X-Groups is the actual implementation of members’ suggestions into their environment as an interaction into an immediate space of use.
Read more3D Head Pose Estimation Based on Scene Flow and Generic Head Model
Head pose is an important indicator of a person's attention, gestures, and communicative behavior with applications in human computer interaction, multimedia and vision systems. In this paper, we present a novel head pose estimation system by performing head region detection using the Kinect [2], followed by face detection, feature tracking, and finally head pose estimation using an active camera. Ten feature points on the face are defined and tracked by an Active Appearance Model (AAM). We propose to use the scene flow approach to estimate the head pose from 2D video sequences. This estimation is based upon a generic 3D head model through the prior knowledge of the head shape and the geometric relationship between the 2D images and a 3D generic model. We have tested our head pose estimation algorithm with various cameras at various distances in real time. The experiments demonstrate the feasibility and advantages of our system.
Read moreDAVE: Detecting Agitated Vocal Events.
DAVE is a comprehensive set of event detection techniques to monitor and detect 5 important verbal agitations: asking for help, verbal sexual advances, questions, cursing, and talking with repetitive sentences. The novelty of DAVE includes combining acoustic signal processing with three different text mining paradigms to detect verbal events (asking for help, verbal sexual advances, and questions) which need both lexical content and acoustic variations to produce accurate results. To detect cursing and talking with repetitive sentences we extend word sense disambiguation and sequential pattern mining algorithms. The solutions have applicability to monitoring dementia patients, for online video sharing applications, human computer interaction (HCI) systems, home safety, and other health care applications. A comprehensive performance evaluation across multiple domains includes audio clips collected from 34 real dementia patients, audio data from controlled environments, movies and Youtube clips, online data repositories, and healthy residents in real homes. The results show significant improvement over baselines and high accuracy for all 5 vocal events.
Read moreA Systematic Technical Review and Architecture of Smart Power Wheelchair Manoeuvring for People with Disabilities and the Elderly: A perspective from Human Computer Interaction and Shared Control Systems
This paper presents a systematic technical review of the human-computer interaction (HCI), Shared Control Systems (SCS), and pervasive System of Systems (SoS) for safe indoor maneuvering in a closed environments using Smart Power Wheelchairs designed for individuals with limited abilities and the elderly. Techniques from HCI, SCS and SoS employed by different researchers in the development of smart wheelchairs are studied and examined. Their advantages and drawbacks in terms of development and ease of operation are discussed and tabulated. Finally, an exemplary architecture for the HCI and SCS of pervasive SoS technologies is proposed, along with key components and shared control system models for the design and development of smart power wheelchairs.
Read moreCNN-Based Facial Expression Recognition from Annotated RGB-D Images for Human–Robot Interaction
Facial expression recognition has been widely used in human computer interaction (HCI) systems. Over the years, researchers have proposed different feature descriptors, implemented different classification methods, and carried out a number of experiments on various datasets for automatic facial expression recognition. However, most of them used 2D static images or 2D video sequences for the recognition task. The main limitations of 2D-based analysis are problems associated with variations in pose and illumination, which reduce the recognition accuracy. Therefore, an alternative way is to incorporate depth information acquired by 3D sensor, because it is invariant in both pose and illumination. In this paper, we present a two-stream convolutional neural network (CNN)-based facial expression recognition system and test it on our own RGB-D facial expression dataset collected by Microsoft Kinect for XBOX in unspontaneous scenarios since Kinect is an inexpensive and portable device to capture both RGB and depth information. Our fully annotated dataset includes seven expressions (i.e., neutral, sadness, disgust, fear, happiness, anger, and surprise) for 15 subjects (9 males and 6 females) aged from 20 to 25. The two individual CNNs are identical in architecture but do not share parameters. To combine the detection results produced by these two CNNs, we propose the late fusion approach. The experimental results demonstrate that the proposed two-stream network using RGB-D images is superior to that of using only RGB images or depth images.
Read more