ENGLISH

The Use of Recurrent Neural Networks in Continuous Speech Recognition

Book information

Language
english
Format
PDF
Filesize
332 kB (339593 bytes)
Pages
\46
Library
twirpx
Time added
2017-08-07 07:01:42

Description

Издательство Cambridge University Engineering Department, 1995. -46 pp.Most - if not all - automatic speech recognition systems explicitly or implicitly compute a score (equivalently, distance, probability, etc.) indicating how well a novel utterance matches a model of the hypothesised utterance. A fundamental problem in speech recognition is how this score may be computed, given that speech is a non-stationary stochastic process. In the interest of reducing the computational complexity, the standard approach used in the most prevalent systems (e.g., dynamic time warping (DTW) and hidden Markov models (HMMs)) factors the hypothesis score into a local acoustic score and a local transition score. In the HMM framework, the observation term models the local (in time) acoustic signal as a stationary process, while the transition probabilities are used to account for the time-varying nature of speech.This chapter presents an extension to the standard HMM framework which addresses the issue of the observation probability computation. Specifically, an artificial recurrent neural network (RNN) is used to compute the observation probabilities within the HMM framework. This provides two enhancements to standard HMMs; (1) the observation model is no longer local, and (2) the RNN architecture provides a nonparametric model of the acoustic signal. The result is a speech recognition system able to model long-term acoustic context without strong assumptions on the distribution of the observations. One such system has been successfully applied to a 20,000 word, speaker-independent, continuous speech recognition task and is described in this chapter.IntroductionThe Hybrid RNN/HMM Approach The HMM Framework Context Modelling Recurrent Networks for Phone Probability EstimationSystem Description The Acoustic Vector Level The Phone Probability Level Posterior Probabilities to Scaled Likelihoods Decoding Scaled LikelihoodsSystem Training Training the RNN RNN Objective Function Gradient Computation Weight UpdateSpecial Features Connectionist Model Combination Duration Modelling Efficient Models Decoding Search Algorithm PruningSummary of Variations Training Criterion Distribution Assumptions Practical IssuesA Large Vocabulary System

Similar books