• Home
  • Search
  • Introducing Self-Supervised Learning Models for Spoken Query-Spoken Term Detection
  • https://doi.org/10.1109/apsipaasc65261.2025.11249196Copy DOI Icon

Introducing Self-Supervised Learning Models for Spoken Query-Spoken Term Detection

  • Oct 22, 2025
  • M Nagase +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

This study was undertaken to improve the performance of Spoken Query Spoken Term Detection (SQ-STD) by introducing Self-Supervised Learning (SSL) models. To construct posteriorgrams, the SSL models are augmented with a Connectionist Temporal Classification (CTC) layer and extracted posterior probability vectors, which represent phonemes or syllables from the layer immediately preceding the CTC output layer. Posteriorgram matching is performed using posteriorgrams generated from SSL models of three types known for their high speech recognition accuracy: wav2vec 2.0, HuBERT, and WavLM. Experiment results obtained on the NTCIR evaluation sets demonstrate that our approach achieves retrieval accuracy exceeding 92 %, which is the state of the art for the evolution sets.

Similar Papers
  • Research Article
  • Citations36

Advancing Acoustic-to-Word CTC Model With Attention and Mixed-Units

  • Sep 04, 2019
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Amit Das +4
  • Conference Article
  • Citations25

Joint Masked CPC And CTC Training For ASR

  • Jun 06, 2021
  • Chaitanya Talnikar +3
  • Research Article

Self-Supervised Deep Learning Models for Low-Resource NLP Applications

  • Feb 26, 2026
  • International Journal of Research and Review in Applied Science, Humanities, and Technology
  • Pardeep Kaur
  • Conference Article
  • Citations7

Investigating Sequence-Level Normalisation For CTC-Like End-to-End ASR

  • May 23, 2022
  • Zeyu Zhao +1
  • Conference Article
  • Citations41

Acoustic-to-word model without OOV

  • Dec 01, 2017
  • 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
  • Jinyu Li +4
  • Conference Article
  • Citations1

K2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning

  • Jun 30, 2025
  • Yifan Yang +11
  • Research Article
  • Citations151

GTC: Guided Training of CTC towards Efficient and Accurate Scene Text Recognition

  • Apr 03, 2020
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Wenyang Hu +4
  • PDF
  • Research Article
  • Citations35

CheSS: Chest X-Ray Pre-trained Model via Self-supervised Contrastive Learning

  • Jan 26, 2023
  • Journal of Digital Imaging
  • Kyungjin Cho +11
  • Conference Article
  • Citations24

Advancing Multi-Accented Lstm-CTC Speech Recognition Using a Domain Specific Student-Teacher Learning Paradigm

  • Dec 01, 2018
  • Shahram Ghorbani +2
  • Conference Article
  • Citations6

Regularizing CTC in Expectation-Maximization Framework with Application to Handwritten Text Recognition

  • Jul 18, 2021
  • Likun Gao +2
  • Book Chapter
  • Citations9

Handwritten Text Recognition with Convolutional Prototype Network and Most Aligned Frame Based CTC Training

  • Jan 01, 2021
  • Likun Gao +2
  • PDF
  • Research Article
  • Citations21

Accented Speech Recognition Based on End-to-End Domain Adversarial Training of Neural Networks

  • Sep 10, 2021
  • Applied Sciences
  • Hyeong-Ju Na +1
  • Conference Article

Voice Command Accuracy Improvement Using Connectionist Temporal Classification in Speech-to-Text Interfaces

  • Dec 01, 2025
  • Shokhista Makhmaraimova +4
  • Conference Article
  • Citations15

Comparing CTC and LFMMI for Out-of-Domain Adaptation of wav2vec 2.0 Acoustic Model

  • Aug 30, 2021
  • Apoorv Vyas +2
  • Conference Article
  • Citations19

Personalization of CTC Speech Recognition Models

  • Jan 09, 2023
  • Saket Dingliwal +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.