Software-hardware co-design: towards ultimate efficiency in deep learning acceleration
Machine learning has surged in popularity in recent years, with deep neural networks (DNNs) becoming a cornerstone of applications such as autonomous driving, augmented reality, and natural language processing, owing to their remarkable accuracy and scalability. As the demand for greater portability and flexibility grows, there is a strong push to optimize these applications for resource-constrained edge devices like smartphones and IoT devices.However, deploying DNN models on such devices presents significant challenges, primarily due to the large size of these models, which is a key contributor to their high performance. For instance, the DeepSpeech 2 model outperforms its predecessor, DeepSpeech 1, but its size is eight times larger, encompassing approximately 68 million parameters. Similarly, other networks designed for various tasks, including convolutional neural networks like VGG and transformer-based networks like BERT and GPT, also exhibit substantial sizes. For example, the GPT-3 model, a predecessor to ChatGPT, consists of an astonishing 175 billion parameters. Consequently, fitting DNN models onto resource-limited edge devices is a formidable challenge. My thesis addresses these challenges with a three-part approach: * High-Performance DNN Accelerators: I will develop DNN accelerators that excel in high performance, low latency, and energy efficiency through an algorithm and hardware co-design across diverse FPGA architectures (Chapter 2). * Real-Time DNN Execution: I aim to achieve exceptional accuracy and real-time execution for various large-scale DNNs on off-the-shelf mobile devices, IoT devices, and other platforms (Chapter 3). * CompVQC Framework: I will explore a systematic framework, CompVQC, to reduce the circuit length of quantum neural networks (QNNs) through model compression, which can achieve up to a 20% improvement in accuracy. This reveals that compilation-aware compression can enhance the robustness of QNNs on near-term noisy quantum devices (Chapter 4). The field of AI has recently undergone a significant transformation, ushering in the era of generative AI, where models like ChatGPT and Stable Diffusion create realistic and creative content for a wide range of applications. As generative AI becomes more prevalent, it is crucial to adapt algorithm-hardware co-design strategies to meet the specific needs of these models, particularly on edge devices. In the final chapter (Chapter 5), I will introduce my ongoing and future work focused on activating generative AI on the edge.--Author's abstract
Read more