Data processing: speech signal processing – linguistics – language – Speech signal processing – Recognition
Reexamination Certificate
1997-03-14
2001-07-10
Tsang, Fan (Department: 2645)
Data processing: speech signal processing, linguistics, language
Speech signal processing
Recognition
C704S241000
Reexamination Certificate
active
06260013
ABSTRACT:
BACKGROUND OF THE INVENTION
The function of automatic speech recognition (ASR) systems is to determine the lexical identity of spoken utterances. The recognition process, also referred to as classification, typically begins with the conversion of an analog acoustical signal into a stream of digitally represented spectral vectors or frames which describe important characteristics of the signal at successive time intervals. The classification or recognition process is based upon the availability of reference models which describe aspects of the behavior of spectral frames corresponding to different words. A wide variety of models have been developed but they all share the property that they describe the temporal characteristics of spectra typical to particular words or sub- word segments. The sequence of spectral vectors arising from an input utterance is compared with the models and the success with which models of different words predict the behavior of the input frames, determines the putative identity of the utterance.
Currently most systems utilize some variant of a statistical model called the Hidden Markov Model (HMM). Such models consist of sequences of states connected by arcs, and a probability density function (pdf) associated with each state describes the likelihood of observing any given spectral vector at that state. A separate set of probabilities may be provided which determine transitions between states.
The process of computing the probability that an unknown input utterance corresponds to a given model, also known as decoding, is usually done in one of two standard ways. The first approach is known as the Forward-Backward algorithm, and uses an efficient recursion to compute the match probability as the sum of the probabilities of all possible alignments of the input sequence and the model states permitted by the model topology. An alternative, called the Viterbi algorithm, approximates the summed match probability by finding the single sequence of model states with the maximum probability. The Viterbi algorithm can be viewed as simultaneously performing an alignment of the input utterance and the model and computing the probability of that alignment.
HMMs can be created to model entire words, or alternatively, a variety of sub-word linguistic units, such as phonemes or syllables. Phone-level HMMs have the advantage that a relatively compact set of models can be used to build arbitrary new words, given that their phonetic transcription is known. More sophisticated versions reflect the fact that contextual effects can cause large variations in the way different phones are realized. Such models are known as allophonic or context-dependent. A common approach is to initiate the search with relatively inexpensive context-independent models and re-evaluate a small number of promising candidates with context-dependent phonetic models.
As in the case of the phonetic models, various levels of modeling power are available in the case of the probability densities describing the observed spectra associated with the states of the HMM. There are two major approaches: the discrete pdf and the continuous pdf. In the former, the spectral vectors corresponding to the input speech are first quantized with a vector quantizer which assigns each input frame an index corresponding to the closest vector from a codebook of prototypes. Given this encoding of the input, the pdfs take on the form of vectors of probabilities, where each component represents the probability of observing a particular prototype vector given a particular HMM state. One of the advantages of this approach is that it makes no assumptions about the nature of such pdfs, but this is offset by the information loss incurred in the quantization stage.
The use of continuous pdfs eliminates the quantization step, and the probability vectors are replaced by parametric functions which specify the probability of any arbitrary input spectral vector given a state. The most common class of functions used for this purpose is the mixture of Gaussians, where arbitrary pdfs are modeled by a weighted sum of Normal distributions. One drawback of using continuous pdfs is that, unlike in the case of the discrete pdf, the designer must make explicit assumptions about the nature of the pdf being modeled—something which can be quite difficult since the true distribution form for the speech signal is not known. In addition, continuous pdf models are computationally far more expensive than discrete pdf models, since following vector quantization the computation of a discrete probability involves no more than a single table lookup.
The probability values in the discrete pdf case and the parameter values of the continuous pdf are most commonly trained using the Maximum Likelihood method. In this manner, the model parameters are adjusted so that the likelihood of observing the training data given the model is maximized. However, it is known that this approach does not necessarily lead to the best recognition performance and this realization has led to the development of new training criteria, known as discriminative, the objective of which is to adjust model parameters so as to minimize the number of recognition errors rather than fit the distributions to the data.
As used heretofore, discriminative training has been applied most successfully to small-vocabulary tasks. In addition, it presents a number of new problems, such as how to appropriately smooth the discriminatively-trained pdfs and how to adapt these systems to a new user with a relatively small amount of training data.
To achieve high recognition accuracies, a recognition system should use high-resolution models which are computationally expensive (e.g., context-dependent, discriminatively-trained continuous density models). In order to achieve real-time recognition, a variety of speedup techniques are usually used.
In one typical approach, the vocabulary search is performed in multiple stages or passes, where each successive pass makes use of increasingly detailed and expensive models, applied to increasingly small lists of candidate models. For example, context independent, discrete models can be used first, followed by context-dependent continuous density models. When multiple sets of models are used sequentially during the search, a separate simultaneous alignment and pdf evaluation is essentially carried out for each set.
In other prior art approaches, computational speedups are applied to the evaluation of the high-resolution pdfs. For example, Gaussian-mixture models are evaluated by a fast but approximate identification of those mixture components which are most likely to make a significant contribution to the probability and a subsequent evaluation of those components in full. Another approach speeds up the evaluation of Gaussian-mixture models by exploiting a geometric approximation of the computation. However, even with speedups the evaluation can be slow enough that only a small number can be carried out.
In another scheme, approximate models are first used to compute the state probabilities given the input speech. All state probabilities which exceed some threshold are then recomputed using the detailed model, the rest are retained as they are. Given the new, composite set of probabilities a new Viterbi search is performed to determine the optimal alignment and overall probability. In this method, the alignment has to be repeated, and in addition, the approximate and detailed probabilities must be similar, compatible quantities. If the detailed model generates probabilities which are significantly higher than those from the approximate models the combination of the two will most likely not lead to satisfactory performance. This requirement constrains this method to use approximate and detailed models which are fairly closely related and thus generate probabilities of comparable magnitude. It should also be noted that in this method there is no guarantee that all of the individual state probabilities that make up the final alignment probability come from detailed mode
Bromberg & Sunstein LLP
Lernout & Hauspie Speech Products N.V.
Opsasnick Michael N.
Tsang Fan
LandOfFree
Speech recognition system employing discriminatively trained... does not yet have a rating. At this time, there are no reviews or comments for this patent.
If you have personal experience with Speech recognition system employing discriminatively trained..., we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Speech recognition system employing discriminatively trained... will most certainly appreciate the feedback.
Profile ID: LFUS-PAI-O-2486721