- Research Article
28
- 10.1016/j.csl.2019.101052
Sequence labeling to detect stuttering events in read speech
- Dec 04, 2019
- Computer Speech & Language
- Sadeen Alharbi + 4 more +4
Sequence labeling to detect stuttering events in read speech
In order to improve the performance of a deep-learning neural network, the paper outlines a stack-based approach incorporating various information sources. A named entity recognition system for Amharic was implemented using a recurrent neural network, a bi-directional long short term memory model. Word vectors based on semantic information were built using an unsupervised learning algorithm, word2vec, while a Conditional Random Fields (CRF) classifier was trained on language independent features to predict each token’s named entity class. The predictions, features and word vectors were fed to the deep neural network to assign labels to the words. This stack-based approach reached an 74.26% F-score, outperforming various other deep-learning set-ups, as well as a baseline CRF classifier, and an ensemble method incorporating the same information sources.
Sequence labeling to detect stuttering events in read speech
Sequence labeling to detect stuttering events in read speech
The Sixth International Symposium on Neural Networks (ISNN 2009)
The Sixth International Symposium on Neural Networks (ISNN 2009)
Comparative Study of Mortality Rate Prediction Using Data-Driven Recurrent Neural Networks and the Lee–Carter Model
The Lee–Carter model could be considered as one of the most important mortality prediction models among stochastic models in the field of mortality. With the recent developments of machine learning and deep learning, many studies have applied deep learning approaches to time series mortality rate predictions, but most of them only focus on a comparison between the Long Short-Term Memory and the traditional models. In this study, three different recurrent neural networks, Long Short-Term Memory, Bidirectional Long Short-Term Memory, and Gated Recurrent Unit, are proposed for the task of mortality rate prediction. Different from the standard country level mortality rate comparison, this study compares the three deep learning models and the classic Lee–Carter model on nine divisions’ yearly mortality data by gender from 1966 to 2015 in the United States. With the out-of-sample testing, we found that the Gated Recurrent Unit model showed better average MAE and RMSE values than the Lee–Carter model on 72.2% (13/18) and 67.7% (12/18) of the database, respectively, while the same measure for the Long Short-Term Memory model and Bidirectional Long Short-Term Memory model are 50%/38.9% (MAE/RMSE) and 61.1%/61.1% (MAE/RMSE), respectively. If we consider forecasting accuracy, computing expense, and interpretability, the Lee–Carter model with ARIMA exhibits the best overall performance, but the recurrent neural networks could also be good candidates for mortality forecasting for divisions in the United States.
Read moreConditional Random Fields as Recurrent Neural Networks
Pixel-level labelling tasks, such as semantic segmentation, play a central role in image understanding. Recent approaches have attempted to harness the capabilities of deep learning techniques for image recognition to tackle pixel-level labelling tasks. One central issue in this methodology is the limited capacity of deep learning techniques to delineate visual objects. To solve this problem, we introduce a new form of convolutional neural network that combines the strengths of Convolutional Neural Networks (CNNs) and Conditional Random Fields (CRFs)-based probabilistic graphical modelling. To this end, we formulate mean-field approximate inference for the Conditional Random Fields with Gaussian pairwise potentials as Recurrent Neural Networks. This network, called CRF-RNN, is then plugged in as a part of a CNN to obtain a deep network that has desirable properties of both CNNs and CRFs. Importantly, our system fully integrates CRF modelling with CNNs, making it possible to train the whole deep network end-to-end with the usual back-propagation algorithm, avoiding offline post-processing methods for object delineation. We apply the proposed method to the problem of semantic image segmentation, obtaining top results on the challenging Pascal VOC 2012 segmentation benchmark.
Read moreTowards subject independent continuous sign language recognition: A segment and merge approach
Towards subject independent continuous sign language recognition: A segment and merge approach
A Recurrent Neural Fuzzy Network
Besides the feedforward neural networks, there are the recurrent networks, where the impulses can be transmitted in both directions due to some reaction connections in these networks. Recurrent Neural Networks (RNNs) are linear or nonlinear dynamic systems. The dynamic behavior presented by the recurrent neural networks can be described both in continuous time, by differential equations and at discrete times by the recurrence relations (difference equations). The distinction between recurrent (or dynamic) neural networks and static neural networks is due to recurrent connections both between the layers of neurons of these networks and within the same layer, too. The aim of this chapter is to describe a Recurrent Fuzzy Neural Network (RFNN) model, whose learning algorithm is based on the Improved Particle Swarm Optimization (IPSO) method.
Read moreDeep Bi-directional Long Short-Term Memory Model for Short-Term Traffic Flow Prediction
Short-term traffic flow prediction plays an important role in intelligent transportation system. Numerous researchers have paid much attention to it in the past decades. However, the performance of traditional traffic flow prediction methods is not satisfactory, for those methods cannot describe the complicated nonlinearity and uncertainty of the traffic flow precisely. Neural networks were used to deal with the issues, but most of them failed to capture the deep features of traffic flow and be sensitive enough to the time-aware traffic flow data. In this paper, we propose a deep bi-directional long short-term memory (DBL) model by introducing long short-term memory (LSTM) recurrent neural network, residual connections, deeply hierarchical networks and bi-directional traffic flow. The proposed model is able to capture the deep features of traffic flow and take full advantage of time-aware traffic flow data. Additionally, we introduce the DBL model, regression layer and dropout training method into a traffic flow prediction architecture. We evaluate the prediction architecture on the dataset from Caltrans Performance Measurement System (PeMS). The experiment results demonstrate that the proposed model for short-term traffic flow prediction obtains high accuracy and generalizes well compared with other models.
Read moreThe Multi-Recurrent Neural Network for State-Of-The-Art Time-Series Processing
The Multi-Recurrent Neural Network for State-Of-The-Art Time-Series Processing
ULISBOA at SemEval-2017 Task 12: Extraction and classification of temporal expressions and events
This paper presents our approach to participate in the SemEval 2017 Task 12: Clinical TempEval challenge, specifically in the event and time expressions span and attribute identification subtasks (ES, EA, TS, TA). Our approach consisted in training Conditional Random Fields (CRF) classifiers using the provided annotations, and in creating manually curated rules to classify the attributes of each event and time expression. We used a set of common features for the event and time CRF classifiers, and a set of features specific to each type of entity, based on domain knowledge. Training only on the source domain data, our best F-scores were 0.683 and 0.485 for event and time span identification subtasks. When adding target domain annotations to the training data, the best F-scores obtained were 0.729 and 0.554, for the same subtasks. We obtained the second highest F-score of the challenge on the event polarity subtask (0.708). The source code of our system, Clinical Timeline Annotation (CiTA), is available at https://github.com/lasigeBioTM/CiTA.
Read moreMedical Named Entity Extraction from Chinese Resident Admit Notes Using Character and Word Attention-Enhanced Neural Network.
The resident admit notes (RANs) in electronic medical records (EMRs) is first-hand information to study the patient’s condition. Medical entity extraction of RANs is an important task to get disease information for medical decision-making. For Chinese electronic medical records, each medical entity contains not only word information but also rich character information. Effective combination of words and characters is very important for medical entity extraction. We propose a medical entity recognition model based on a character and word attention-enhanced (CWAE) neural network for Chinese RANs. In our model, word embeddings and character-based embeddings are obtained through character-enhanced word embedding (CWE) model and Convolutional Neural Network (CNN) model. Then attention mechanism combines the character-based embeddings and word embeddings together, which significantly improves the expression ability of words. The new word embeddings obtained by the attention mechanism are taken as the input to bidirectional long short-term memory (BI-LSTM) and conditional random field (CRF) to extract entities. We extracted nine types of key medical entities from Chinese RANs and evaluated our model. The proposed method was compared with two traditional machine learning methods CRF, support vector machine (SVM), and the related deep learning models. The result shows that our model has better performance, and the result of our model reaches 94.44% in the F1-score.
Read moreRecurrent neural networks: methods and applications to non-linear predictions
This thesis deals with recurrent neural networks, a particular class of artificial neural networks which can learn a generative model of input sequences. The input is mapped, through a feedback loop and a non-linear activation function, into a hidden state, which is then projected into the output space, obtaining either a probability distribution or the new input for the next time-step. This work consists mainly of two parts: a theoretical study for helping the understanding of recurrent neural networks framework, which is not yet deeply investigated, and their application to non-linear prediction problems, since recurrent neural networks are really powerful models suitable for solving several practical tasks in different fields. For what concerns the theoretical part, we analyse the weaknesses of state-of-the-art models and tackle them in order to improve the performance of a recurrent neural network. Firstly, we contribute in the understanding of the dynamical properties of a recurrent neural network, highlighting the close relation between the definition of stable limit cycles and the echo state property of an echo state network. We provide sufficient conditions for the convergence of the hidden state to a trajectory, which is uniquely determined by the input signal, independently of the initial states. This may help extend the memory of the network and increase the design options for the network. Moreover, we develop a novel approach to address the main problem in training recurrent neural networks, the so-called vanishing gradient problem. Our new method allows us to train a very simple recurrent neural network, making the gradient not to vanish even after many time-steps. Exploiting the singular value decomposition of the vanishing factors in the gradient and random matrices theory, we find that the singular values have to be confined in a narrow interval and derive conditions about their root mean square value. Then, we also improve the efficiency of the training of a recurrent neural network, defining a new method for speeding up this process. Thanks to a least square regularization, we can initialize the parameters of the network, in order to set them closer to the minimum and running fewer epochs of classical training algorithms. Moreover, it is also possible to completely train the network with our initialization method, running more iterations of it without losing in performance with respect to classical training algorithms. Finally, it is also possible to use it as a real-time learning algorithm, adjusting the parameters to the new data through one iteration of our initialization. In the last part of this thesis, we apply recurrent neural networks to non-linear prediction problems. We consider prediction of numerical sequences, estimating the following input choosing it from a probability distribution. We study an automatic text generation problem, where we need to predict the following character in order to compose words and sentences, and a path prediction of walking mobile users in the central area of a city, as a sequence of crossroads. Then, we analyse the prediction of video frames, discovering a wide range of applications related to the prediction of movements. We study the collision problem of bouncing balls, taking into account only the sequence of video frames without any knowledge about the physical characteristics of the problem, and the distribution over days of mobile user in a city and in a whole region. Finally, we address the state-of-the-art problem of missing data imputation, analysing the incomplete spectrogram of audio signals. We restore audio signals with missing time-frequency data, demonstrating via numerical experiments that a performance improvement can be achieved involving recurrent neural networks.
Read moreClassification of Covid-19 Tweets Using Deep Learning Techniques
In this digital era, there is an exponential growth of text-based content in the electronic world. Data as texts exist in the form of documents, social media posts on Facebook, Twitter, etc., logs, sensor data, and emails. Twitter is a social platform where users express their views on various aspects in a day to day life. Twitter produces over 500 million tweets daily that is 6000 tweets per second. Twitter data is, by definition, very noisy and unstructured in nature. Text classifications based on the machine learning techniques have problems like poor generalization ability and sparsity dimension explosion. Classifiers based on deep learning techniques are implemented to improve accuracy to overcome shortcomings of machine learning techniques and to avoid feature extraction processes and have high prediction accuracy and strong learning ability. In this work, the classification of tweets is performed on Covid-19 dataset by implementing deep learning techniques namely Convolution Neural Network (CNN), Recurrent Neural Network (RNN), Recurrent Convolution Neural Network (RCNN), Recurrent Neural Network with Long Short Term Memory (RNN+LSTM), and Bidirectional Long Short Term Memory with Attention (BI-LSTM + Attention). The algorithms are implemented using two-word embedding techniques namely Global Vectors for Word Representation (GloVe) and Word2Vec. RNN with Bidirectional LSTM model has performed better than all the classifiers considered. It has classified the text with an accuracy of 93% and above when used with GloVe and Word2Vec.
Read moreMost Possible Likely Word Generation for Text based Applications using Generative Pretrained Transformer model Comparing to Long Short Term Memory Model
In contrast to long short term memory techniques, the proposed study intends to produce automatic next word generation for text-based applications while enhancing accuracy using state-of-the-art generative pretrained transformers and recurrent neural networks. Materials and Methods: On the data, which is a text file including a series of words, generative pretrained transformer models and long short term memory are used. Long short term memory model that contrasts cutting edge generative pretrained transformer models for recommendation accuracy of the next word. It has been suggested and created to use LSTM. The sample size was calculated to be 8046 for each group with a G power of 0.8. Results: When compared with a long short term memory model (70.84%) for the same dataset p=0.02(p<0.05), the accuracy in predicting the upcoming word for text editor based Applications utilizing generative pretrained transformers was greatest at 87.98% with the lowest mean error. In generating the next word for text-based Applications, the study shows that generative pretrained transformers are more accurate than long short term memory..Conclusion: The study demonstrates that when recommending the most possible word for text generation based applications, generative pretrained transformers are more accurate than long short term memory.
Read moreEmbeddings in Natural Language Processing: Theory and Advances in Vector Representations of Meaning
Embeddings in Natural Language Processing: Theory and Advances in Vector Representations of Meaning
A Novel Modular Recurrent Wavelet Neural Network and Its Application to Nonlinear System Identification
To reduce the computational complexity and improve the performance of the recurrent wavelet neural network (RWNN), a novel modular recurrent neural network based on the pipelined architecture (PRWNN) with low computational complexity is presented in this paper. Its modified adaptive real-time recurrent learning (RTRL) algorithm is derived on the gradient descent approach. The PRWNN comprises a number of RWNN modules that are cascaded in a chained form and inherits the modular architectures of the pipelined recurrent neural network (PRNN) proposed by Haykin and Li. Since those modules of the PRWNN can be performed simultaneously in a pipelined parallelism fashion, it would result in a significant improvement in computational efficiency. And the performance of the PRWNN can be also further improved. Computer simulations have demonstrated that the PRWNN provides considerably better performance compared to the single RWNN model for nonlinear dynamic system identification.
Read more