Speech trigger and speech command recognition technologies have become pivotal human-machine interfaces for AIoT applications. Typically, a speech-related system consists of a mel-filter-based handcrafted feature extractor and an artificial neuron-based classifier. Recent advancements [1–3] have demonstrated promising results by training the feature extractor jointly with the neural network classifier. Moreover, the emergence of spiking neuron network training algorithms and hardware designs has showcased their competitive edge over traditional artificial neural networks. Consequently, the design and training of learnable audio feature extractors and spiking neural network classifiers are the key focuses of this thesis. The first focus of this thesis is the design of an ultra-low-power, learnable audio feature extractor to support end-to-end training. Conventional low-power approaches, such as extreme pruning and quantization in the audio feature extractor (AuFEx) design, often result in decreased accuracy. To address this, the thesis proposes a hardware-aware, end-to-end training framework to co-optimize the AuFEx design, achieving high performance with minimal hardware cost. Two types of AuFEx designs are explored: a time-domain (TD) filter-based learnable audio feature extractor (AuFEx-TD) and a frequency-domain (FD) filter-based learnable audio feature extractor (AuFEx-FD). Using this framework, the AuFEx-TD achieves high accuracy (97.8%) for one wake word detection, with low power consumption (496 nW) and low latency (375 μs). The optimized AuFEx-FD demonstrates noise-robust 10-keyword spotting accuracy (87.5%@5dB – 92.2%@20dB) with low power consumption (1.26 μW). Both AuFEx-TD and AuFEx-FD designs are implemented using UMC 40nm CMOS technology to validate the framework's effectiveness. For the speech trigger or wake word detection (WWD) task, power and latency are critical design considerations. This thesis presents a sub-μW spiking neural network accelerator with an 8 ms latency. The core WWD engine employs a spiking convolutional neural network (SCNN) model, leveraging sparse activation and addition-only operations within spiking neurons. The proposed SCNN model enhances the existing framewise incremental computation flow [4] by incorporating a spike processing unit (SPU), reducing system power and latency by 16.5% and 43.2%, respectively. Extensive quantization further reduces weight and activation precisions to 4-bit and 1-bit. Additionally, a power-gating mechanism, driven by an energy-based voice activity detection (VAD) module, further minimizes power consumption in random and sparse event (RSE) scenarios. Full chip simulation results show that the chip consumes only 110 nW, with a 2.15% false alarm rate and a 3.00% false reject rate in a 10% voice event stream test. It achieves state-of-the-art recognition accuracy of 99% and 96% for one and two wake word detection tasks, respectively. For speech command recognition or keyword spotting (KWS) tasks, robustness against noise and speaker accents is essential for user experience. This thesis explores both end-to-end training and on-device transfer learning to enhance KWS chip robustness with minimal power overhead. The proposed chip comprises three main blocks: the Learnable Audio Feature Extractor (AuFEx), the Spiking-DSCNN Core (SC), and the Transfer Learning Core (TLC). The AuFEx is an optimized AuFEx-FD, with the on-chip memory for storing FD-filter parameters replaced by an FD-filter parameter generator, reducing power consumption to 1.225 μW. The SC implements a spiking neural network (SNN) model using a Depthwise Separable Convolutional Neural Network (DSCNN) topology. Multi-level encoding in spiking neurons reduces simulation time, while the inherent activation sparsity of these neurons decreases power consumption and latency. The TLC supports both forward and backward propagation for a fully connected layer and softmax layer, enabling the KWS chip to adapt to environmental and user variations using minimal labeled data and weight updates. Fabricated in UMC 40nm CMOS technology, the chip achieves 10-keyword spotting accuracy ranging from 87.5% to 92.2% under background noise levels of 5dB to 20dB. Through on-device transfer learning, the chip can improve KWS accuracy on personalized datasets by 11.73%, using only a few examples per keyword and limited training epochs. The simulated power consumption of the chip in streaming inference mode is 10.26 μW at a 250 kHz clock frequency and 0.85V supply voltage. In conclusion, this thesis presents an end-to-end spiking neural network-based speech system for speech trigger (WWD) and speech command recognition (KWS) tasks. By integrating the AuFEx-TD and the SCNN accelerator, as described in the first and second parts of the thesis, a low-latency, sub-μW wake word detection system has been developed. In the final part, a noise-robust, personalized, few-μW-level speech command recognition ASIC is introduced, demonstrating the feasibility of high-performance, low-power AIoT speech interfaces.
Read more