- Research Article
15
- 10.11834/jig.230020
Human-computer interaction for virtual-real fusion
- Jan 01, 2023
- Journal of Image and Graphics
- Jianhua Tao + 5 more +5
Human-computer interaction for virtual-real fusion
This project focuses on creating a versatile AI desktop assistant using Python and Tkinter, designed to simplify routine tasks through intuitive voice commands. By combining speech recognition, text-to speech technologies, and a graphical user interface, the assistant delivers a seamless user experience for task automation. The project serves as an excellent starting point for those interested in artificial intelligence, human-computer interaction, and Python programming. The assistant's core functionality includes recognizing and executing voice commands, providing real-time weather updates for a specified city, announcing the current time, and automating web browsing for popular websites. A minimalistic yet effective GUI ensures accessibility for users of all skill levels. With its modular architecture, this project also demonstrates how Python's extensive library ecosystem can be leveraged to integrate diverse functionalities efficiently. This project creates a simple AI desktop assistant using Python and Tkinter. The assistant helps automate common tasks through voice commands, using speech recognition and text-to-speech technologies.
Human-computer interaction for virtual-real fusion
Human-computer interaction for virtual-real fusion
Deep Learning for Intelligent Human–Computer Interaction
In recent years, gesture recognition and speech recognition, as important input methods in Human–Computer Interaction (HCI), have been widely used in the field of virtual reality. In particular, with the rapid development of deep learning, artificial intelligence, and other computer technologies, gesture recognition and speech recognition have achieved breakthrough research progress. The search platform used in this work is mainly the Google Academic and literature database Web of Science. According to the keywords related to HCI and deep learning, such as “intelligent HCI”, “speech recognition”, “gesture recognition”, and “natural language processing”, nearly 1000 studies were selected. Then, nearly 500 studies of research methods were selected and 100 studies were finally selected as the research content of this work after five years (2019–2022) of year screening. First, the current situation of the HCI intelligent system is analyzed, the realization of gesture interaction and voice interaction in HCI is summarized, and the advantages brought by deep learning are selected for research. Then, the core concepts of gesture interaction are introduced and the progress of gesture recognition and speech recognition interaction is analyzed. Furthermore, the representative applications of gesture recognition and speech recognition interaction are described. Finally, the current HCI in the direction of natural language processing is investigated. The results show that the combination of intelligent HCI and deep learning is deeply applied in gesture recognition, speech recognition, emotion recognition, and intelligent robot direction. A wide variety of recognition methods were proposed in related research fields and verified by experiments. Compared with interactive methods without deep learning, high recognition accuracy was achieved. In Human–Machine Interfaces (HMIs) with voice support, context plays an important role in improving user interfaces. Whether it is voice search, mobile communication, or children’s speech recognition, HCI combined with deep learning can maintain better robustness. The combination of convolutional neural networks and long short-term memory networks can greatly improve the accuracy and precision of action recognition. Therefore, in the future, the application field of HCI will involve more industries and greater prospects are expected.
Read moreAutomatic bimodal audiovisual speech recognition: A review
Human computer interaction (HCI) is very crucial in our day-to-day activity. Speech is one of the essential and intuitive ways to interact with machines such as Smartphone, which has multiple sensors as microphone, camera, etc. An efficient performance speech recognition system improves interaction between man and machines by making latter more receptive to user needs. Such system has Automatic speech recognition (ASR) engine, which is facing a unique challenge of accuracy in recognition rate. By integrating acoustic signal feature vectors with the visual features, a more robust audiovisual speech recognition engine (AVSR) could be developed for real environmental scenarios. This paper presents past research and development in the field of ASR and AVSR technologies. It describes key technological perspective and admiration of the fundamental progress in ASR and AVSR. The objective of this review is to summarize and compare some of the well-known methods experimented by previous researchers, and to conclude with direction on future research proficiency in HCI system using ASR and AVSR engine.
Read moreA comparative study of voice and graphical user interfaces with respect to literacy levels
Visual and aural are two most important channels of information processing. While most of the interaction with computers have been designed around the visual channel, there are circumstances where voice based man-machine interaction becomes preferable, and in some cases, necessary, given that voice based interaction comes naturally to humans and can be used by illiterate people easily.
Read moreVoice Controlled Desktop Assistant
ABSTRACT: The advancement of Artificial Intelligence (AI) and Natural Language Processing (NLP) has led to the development of voice- controlled systems that enhance Human–Computer Interaction (HCI). This paper presents the design and implementation of a Voice Controlled Desktop Assistant, capable of performing various tasks such as opening applications, searching the web, playing multimedia content, sending messages, and providing real-time system feedback. The assistant uses speech recognition for voice input and text- to-speech synthesis for generating responses. Developed using Python and integrated APIs, it aims to automate desktop functions through intuitive voice commands. The project demonstrates an effective approach for creating user-friendly, hands-free computing environments. Additionally, the system emphasizes lightweight processing to ensure efficient real-time performance without dependency on cloud-based services. It further contributes to the ongoing evolution of intelligent personal assistants designed for offline and secure desktop automation. KEYWORD: Voice Assistant, Artificial Intelligence, Natural Language Processing, Speech Recognition, Automation, Human– Computer Interaction. ABSTRACT: The advancement of Artificial Intelligence (AI) and Natural Language Processing (NLP) has led to the development of voice- controlled systems that enhance Human–Computer Interaction (HCI). This paper presents the design and implementation of a Voice Controlled Desktop Assistant, capable of performing various tasks such as opening applications, searching the web, playing multimedia content, sending messages, and providing real-time system feedback. The assistant uses speech recognition for voice input and text- to-speech synthesis for generating responses. Developed using Python and integrated APIs, it aims to automate desktop functions through intuitive voice commands. The project demonstrates an effective approach for creating user-friendly, hands-free computing environments. Additionally, the system emphasizes lightweight processing to ensure efficient real-time performance without dependency on cloud-based services. It further contributes to the ongoing evolution of intelligent personal assistants designed for offline and secure desktop automation. KEYWORD: Voice Assistant, Artificial Intelligence, Natural Language Processing, Speech Recognition, Automation, Human– Computer Interaction.
Read moreError handling in multimodal voice-enabled interfaces of tour-guide robots using graphical models
Error handling in multimodal voice-enabled interfaces of tour-guide robots using graphical models
PYTHON POWERED INTELLIGENCE AND ML
Python Powered Intelligence And ML is designed to be your essential companion in your journey through the world of Artificial Intelligence and Python programming. We understand the importance of building a solid foundation in AI concepts, as well as mastering the tools and techniques needed to implement AI solutions effectively. What You’ll Find Inside: Foundation of Artificial Intelligence: In Chapter 1, we lay the groundwork for your AI education, providing a strong understanding of the fundamentals. Knowledge Presentation: Chapter 2 delves into how knowledge is represented in AI systems, a crucial element for creating intelligent machines. Informed / Heuristic Search Strategies: Chapter 3 explores strategies for problem-solving and decision-making, crucial in the AI domain. Natural Language Processing: In Chapter 4, we dive into the world of language understanding and processing, a key area of AI. Soft Computing: Chapter 5 introduces the concept of soft computing, which enables AI systems to work with uncertainty and imprecision. Neural Networks: Chapter 6 covers neural networks, a fundamental technology in modern AI, inspired by the human brain. Fuzzy Systems: Chapter 7 is all about fuzzy logic, which allows AI systems to deal with vagueness and uncertainty. History of Genetic Algorithms: Chapter 8 takes you through the fascinating history and principles of genetic algorithms. Regression: In Chapter 9, we explain regression techniques for predictive modeling, a crucial tool in AI. Python Programming: The second part of Python Powered Intelligence And ML focuses on Python, one of the most versatile and popular programming languages. Numpy: Chapter 11 introduces the powerful library for numerical computing in Python. Pandas: In Chapter 12, we explore Pandas, a tool for data manipulation and analysis. Matplotlib: Chapter 13 introduces you to data visualization using Matplotlib. Regression: We revisit regression in Chapter 14, providing more insights and applications. Our aim is to empower you with the knowledge and skills to excel in Artificial Intelligence and Python programming. Whether you are a beginner or an experienced programmer, this book is your go-to resource for AI and Python. We hope you enjoy this journey with us, and may it inspire you to explore the exciting world of Artificial Intelligence and Python programming. Happy learning!
Read moreAn Approach Towards to Real Time AI Desktop Voice Assistant
The advent of artificial intelligence (AI) has revolutionized human-computer interaction, making it more natural and intuitive. This research paper presents the development and implementation of an AI Desktop Voice Assistant designed to enhance productivity and accessibility for users. The voice assistant leverages advanced speech recognition, natural language processing (NLP), and machine learning techniques to understand and execute user commands. Key functionalities include voice-activated application control, web searches, personalized reminders, and real-time information retrieval. Our system integrates with widely-used APIs and services, providing a seamless user experience across various tasks. The paper discusses the architecture of the voice assistant, detailing the integration of components such as the speech recognition engine, NLP models, and the dialogue management system. We explore the challenges encountered during development, including accurate speech recognition in noisy environments and handling ambiguous user commands. Solutions implemented to address these challenges are also presented. Furthermore, we conduct a usability study to evaluate the effectiveness and user satisfaction of the voice assistant. The results indicate a high level of user engagement and satisfaction, demonstrating the practical benefits of incorporating AI-driven voice interfaces into desktop environments. Our findings contribute to the ongoing research in human-computer interaction, suggesting pathways for future enhancements in AI voice assistants..
Read moreNew Speech Noise Reduction Recognition System Based on Spatial Filtering Technology and CI1103 Speech Module
With the continuous development of science and technology in recent years, there are more and more ways of human-computer interaction, and its technology is becoming more and more mature. Among them, the human-computer voice interaction mode occupies an important position, which is gradually integrated into our lives. However, through reviewing and summarizing previous studies, it is found that speech recognition still has some shortcomings in noise reduction processing [2]. Therefore, to design a system that is immune to ambient noise and can perform voice human-computer interaction more accurately, based on spatial filtering technology and CI1103 speech recognition module, a new type of noise reduction speech recognition system using a unique anti-noise processing microphone is proposed. In this new system, combined with the built-in anti-noise filter of CI1103, the noise reduction process during speech recognition is realized with the help of the spatial filtering technology where the new digital signal is used to process IC “BU8332KV-M” [3]. Its peripheral control interfaces such as built-in CPU core and Audio Codec module with high performance and low power, integrated multiple UART, IIC, SPI, PWM, and GPIO can develop all kinds of cost-effective and single-chip intelligent voice products. Experimental results show that the new noise reduction speech recognition system has higher speech recognition accuracy than that of the traditional speech recognition systems in noisy environments, which has the function of rejecting false recognition, and is suitable for various speech systems with environmental noise interference.
Read moreGesture speak: Hands-Free Computer Control with Hand Gestures and Voice Commands
In an era of advancing human-computer interaction, "GestureSpeak" emerges as a pioneering project that facilitates hands-free control over computing devices through intuitive hand gestures and voice commands. Leveraging the power of MediaPipe and OpenCV in Python, this innovative system enables users to seamlessly navigate their digital environments without the need for traditional input devices. By harnessing the capabilities of computer vision and machine learning, GestureSpeak interprets and responds to users' gestures and vocal instructions, opening new frontiers in accessibility and user experience. This abstract offers a glimpse into the transformative potential of GestureSpeak in revolutionizing the way we interact with computers, making technology more accessible and intuitive for all users. Advancements in human-computer interaction have led to the development of innovative projects that redefine how users interact with technology. This paper introduces a novel integration of two projects: GestureSpeak and VoiceRobot, which together enable intuitive and hands-free control over computing devices using hand gestures and voice commands. VoiceRobot complements GestureSpeak by adding voice command capabilities to the interaction model. Powered by speech recognition technology, VoiceRobot enables users to control devices and launch applications using natural language. With VoiceRobot, initiating the GestureSpeak project is as simple as speaking a command. The combined system offers unique benefits, including improved accessibility for users with mobility impairments and a more natural way to interact with computing devices. By integrating gesture recognition and voice commands, the project enhances user engagement and efficiency, paving the way for future advancements in human-computer interaction. Launching the Project with VoiceRobot:A notable feature of this integrated system is the ability to launch the GestureGenie project using VoiceRobot. By speaking a predefined command, such as "Activate GestureGenie," users can initiate the hand gesture control system effortlessly. This capability highlights the seamless integration of gesture recognition and voice command technologies, showcasing the project's versatility and user-friendly design. In conclusion, the integration of GestureGenie and VoiceRobot represents a significant advancement in human-computer interaction, offering a comprehensive solution for hands-free device control. This abstract provides insight into the combined capabilities of these projects and their potential impact on accessibility, user experience, and the future of interactive computing.
Read moreDetection of Face Direction by Implementing Face Edge Patterns
In this paper, a method for detecting the direction of a human face is developed; regardless of its age or sex. The method involves creating a set of five face patterns representing the front, up, down, left, and right directions of a face. The face patterns are produced by applying Canny’s edge detection algorithm on some face files. The direction of the input face is found by first applying the above algorithm on the input file and comparing it with the five face patterns. The face pattern that gives minimum difference will represent the direction of the input face. Excellent results were reported when applied on images with relatively clear background and the head were centered at the image area.
Read moreFifty years of progress in speech and speaker recognition
Speech and speaker recognition technology has made very significant progress in the past 50 years. The progress can be summarized by the following changes: (1) from template matching to corpus-base statistical modeling, e.g., HMM and n-grams, (2) from filter bank/spectral resonance to Cepstral features (Cepstrum + DCepstrum + DDCepstrum), (3) from heuristic time-normalization to DTW/DP matching, (4) from gdistanceh-based to likelihood-based methods, (5) from maximum likelihood to discriminative approach, e.g., MCE/GPD and MMI, (6) from isolated word to continuous speech recognition, (7) from small vocabulary to large vocabulary recognition, (8) from context-independent units to context-dependent units for recognition, (9) from clean speech to noisy/telephone speech recognition, (10) from single speaker to speaker-independent/adaptive recognition, (11) from monologue to dialogue/conversation recognition, (12) from read speech to spontaneous speech recognition, (13) from recognition to understanding, (14) from single-modality (audio signal only) to multi-modal (audio/visual) speech recognition, (15) from hardware recognizer to software recognizer, and (16) from no commercial application to many practical commercial applications. Most of these advances have taken place in both the fields of speech recognition and speaker recognition. The majority of technological changes have been directed toward the purpose of increasing robustness of recognition, including many other additional important techniques not noted above.
Read moreUsing artificial intelligence to automatically test GUI
This position paper presents a synopsis of the imperative function artificial intelligence (AI) has partaken in software engineering (SE) as well as in software testing. In addition, the paper discusses how graphical user interface (GUI), and event driven software testing can derive benefits from the use of AI techniques. Artificial intelligence has significantly aided the process of the automation of different software process. The employment of AI in software testing is not novel, having played a crucial role in the automation of software testing since its innovation. The usage of AI techniques not only reduces the cost but it also guarantees better quality as well as thorough testing. GUI Testing can be considered as the most challenging area of software testing. Although the results are quite preliminary, but the application of different AI techniques for GUI testing has proven to produce ideal results. Nevertheless, the application of AI techniques in GUI testing, in comparison with software testing, which has procured much assistance by venturing with AI, leaves much to be desired.
Read moreAnatomySketch: An Extensible Open-Source Software Platform for Medical Image Analysis Algorithm Development
The development of medical image analysis algorithm is a complex process including the multiple sub-steps of model training, data visualization, human–computer interaction and graphical user interface (GUI) construction. To accelerate the development process, algorithm developers need a software tool to assist with all the sub-steps so that they can focus on the core function implementation. Especially, for the development of deep learning (DL) algorithms, a software tool supporting training data annotation and GUI construction is highly desired. In this work, we constructed AnatomySketch, an extensible open-source software platform with a friendly GUI and a flexible plugin interface for integrating user-developed algorithm modules. Through the plugin interface, algorithm developers can quickly create a GUI-based software prototype for clinical validation. AnatomySketch supports image annotation using the stylus and multi-touch screen. It also provides efficient tools to facilitate the collaboration between human experts and artificial intelligent (AI) algorithms. We demonstrate four exemplar applications including customized MRI image diagnosis, interactive lung lobe segmentation, human-AI collaborated spine disc segmentation and Annotation-by-iterative-Deep-Learning (AID) for DL model training. Using AnatomySketch, the gap between laboratory prototyping and clinical testing is bridged and the development of MIA algorithms is accelerated. The software is opened at https://github.com/DlutMedimgGroup/AnatomySketch-Software.
Read moreMultimodal information fusion method in emotion recognition in the background of artificial intelligence
Recent advances in Semantic IoT data integration have highlighted the importance of multimodal fusion in emotion recognition systems. Human emotions, formed through innate learning and communication, are often revealed through speech and facial expressions. In response, this study proposes a hidden Markov model‐based multimodal fusion emotion detection system, combining speech recognition with facial expressions to enhance emotion recognition rates. The integration of such emotion recognition systems with Semantic IoT data can offer unprecedented insights into human behavior and sentiment analysis, contributing to the advancement of data integration techniques in the context of the Internet of Things. Experimental findings indicate that in single‐modal emotion detection, speech recognition achieves a 76% accuracy rate, while facial expression recognition achieves 78%. However, when state information fusion is applied, the recognition rate increases to 95%, surpassing the national average by 19% and 17% for speech and facial expressions, respectively. This demonstrates the effectiveness of multimodal fusion in emotion recognition, leading to higher recognition rates and reduced workload compared to single‐modal approaches.
Read more