• Home
  • Search
  • Hybrid Transformer/CTC Networks for Hardware Efficient Voice Triggering
  • Open Access IconOpen Access
  • Cite Icon19
  • https://doi.org/10.21437/interspeech.2020-1330Copy DOI Icon

Hybrid Transformer/CTC Networks for Hardware Efficient Voice Triggering

  • Oct 25, 2020
  • Saurabh Adya +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We consider the design of two-pass voice trigger detection systems. We focus on the networks in the second pass that are used to re-score candidate segments obtained from the first-pass. Our baseline is an acoustic model(AM), with BiLSTM layers, trained by minimizing the CTC loss. We replace the BiLSTM layers with self-attention layers. Results on internal evaluation sets show that self-attention networks yield better accuracy while requiring fewer parameters. We add an auto-regressive decoder network on top of the self-attention layers and jointly minimize the CTC loss on the encoder and the cross-entropy loss on the decoder. This design yields further improvements over the baseline. We retrain all the models above in a multi-task learning(MTL) setting, where one branch of a shared network is trained as an AM, while the second branch classifies the whole sequence to be true-trigger or not. Results demonstrate that networks with self-attention layers yield $\sim$60% relative reduction in false reject rates for a given false-alarm rate, while requiring 10% fewer parameters. When trained in the MTL setup, self-attention networks yield further accuracy improvements. On-device measurements show that we observe 70% relative reduction in inference time. Additionally, the proposed network architectures are $\sim$5X faster to train.

Similar Papers
  • Research Article
  • Citations75

MEDIC: a multi-task learning dataset for disaster image classification

  • Sep 03, 2022
  • Neural Computing and Applications
  • Firoj Alam +5
  • Video Transcripts

Annotations Matter: Leveraging Multi-task Learning to Parse UD and SUD

  • Aug 01, 2021
  • Underline Science Inc.
  • Daniel Dakota +1
  • PDF
  • Research Article
  • Citations13

Leveraging Multi-Task Learning to Cope With Poor and Missing Labels of Mammograms.

  • Jan 11, 2022
  • Frontiers in Radiology
  • Mickael Tardy +1
  • Conference Article
  • Citations28

L-Vector: Neural Label Embedding for Domain Adaptation

  • Apr 10, 2020
  • Zhong Meng +6
  • Research Article
  • Citations17

Multi-Task Learning for Analyzing and Sorting Large Databases of Sequential Data

  • Aug 01, 2008
  • IEEE Transactions on Signal Processing
  • Kai Ni +3
  • PDF
  • Research Article
  • Citations16

Personal Interest Attention Graph Neural Networks for Session-Based Recommendation

  • Nov 12, 2021
  • Entropy
  • Xiangde Zhang +3
  • Research Article

Multitask learning via task embeddings for glass property prediction with improved sample efficiency

  • Feb 01, 2026
  • Computational Materials Science
  • Gregor Maier +2
  • PDF
  • Research Article
  • Citations29

Attention-Based Convolution Skip Bidirectional Long Short-Term Memory Network for Speech Emotion Recognition

  • Dec 25, 2020
  • IEEE Access
  • Huiyun Zhang +2
  • PDF
  • Research Article
  • Citations4

A Self-Attention-Based Imputation Technique for Enhancing Tabular Data Quality

  • Jun 04, 2023
  • Data
  • Do-Hoon Lee +1
  • Conference Article
  • Citations327

Anomaly Detection in Video via Self-Supervised and Multi-Task Learning

  • Jun 01, 2021
  • Mariana-Iuliana Georgescu +5
  • Conference Article
  • Citations44

It Takes Nine to Smell a Rat: Neural Multi-Task Learning for Check-Worthiness Prediction

  • Oct 22, 2019
  • Slavena Vasileva +4
  • Conference Article
  • Citations36

Improving Entity Recommendation with Search Log and Multi-Task Learning

  • Jul 01, 2018
  • Jizhou Huang +4
  • Video Transcripts

Unsupervised Multi-Task Domain Adaptation

  • Dec 29, 2020
  • Underline Science Inc.
  • Mei-Chen Yeh
  • Conference Article
  • Citations2

Cross-Stitched Multi-task Dual Recursive Networks for Unified Single Image Deraining and Desnowing

  • Oct 26, 2022
  • Sotiris Karavarsamis +3
  • PDF
  • Research Article
  • Citations37

Multi-task learning for few-shot biomedical relation extraction

  • Apr 19, 2023
  • Artificial Intelligence Review
  • Vincenzo Moscato +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.