Image analysis – Applications
Patent
1993-10-22
1998-06-23
Razavi, Michael T.
Image analysis
Applications
382288, 382291, 382296, G06T 760
Patent
active
057713065
ABSTRACT:
The apparatus for the recognition of speech comprises an acoustic preprocessor, a visual preprocessor, and a speech classifier that operates the acoustic and visual preprocessed data. The acoustic preprocessor comprises a log mel spectrum analyzer that produces an equal mel bandwidth log power spectrum. The visual processor detects the motion of a set of fiducial markers on the speaker's face and extracts a set of normalized distance vectors describing lip and mouth movement. The speech classifier uses a multilevel time-delay neural network operating on the preprocessed acoustic and visual data to form an output probability distribution that indicates the probability of each candidate utterance having been spoken, based on the acoustic and visual data.
REFERENCES:
patent: 4620286 (1986-10-01), Smith et al.
patent: 4706296 (1987-11-01), Pedotti et al.
patent: 4757541 (1988-07-01), Beadles
patent: 4769845 (1988-09-01), Nakamura
patent: 4841575 (1989-06-01), Welsh et al.
patent: 4937872 (1990-06-01), Hopfield et al.
patent: 4975960 (1990-12-01), Petajan
patent: 5022089 (1991-06-01), Wilson
patent: 5163111 (1992-11-01), Baji et al.
patent: 5173945 (1992-12-01), Pieters et al.
Waibel, A., "Modular Construction of Time-Delay Neural Networks for Speech Recognition," Neural Computation 1, pp. 39-46 (1989).
Petajan, E., et al., "An Improved Automatic Lipreading System to Enhance Speech Recognition," ACM SIGCHI-88, pp. 19-25 (1988).
Pentland, A., et al., "Lip Reading:Automatic Visual Recognition of Spoken Words," Proc. Image Understanding and Machine Vision, Optical Society of America, pp. 1-9 (Jun. 12-14, 1989).
Yuhas, B.P., et al., "Integration of Acoustic and Visual Speech Signals Using Neural Networks," IEEE Communications Magazine, pp. 65-71 (Nov. 1989).
Waibel, A., et al., "Phoneme Recognition: Neural Networks vs. Hidden Markov Models," IEEE ICASSP88 Proceedings, vol. 1, pp. 107-110 (1988).
T.J. Sejnowski et al., "Combining Visual and Acoustic Speech Signals with a Neural Network Improves Intelligibility," Advances in Neural Info. Processing Systems 2, 8 pgs. (undated).
Levine Earl Isaac
Stork David G.
Wolff Gregory Joseph
Chang Jon
Razavi Michael T.
Ricoh Company, Ltd
Ricoh Corporation
LandOfFree
Method and apparatus for extracting speech related facial featur does not yet have a rating. At this time, there are no reviews or comments for this patent.
If you have personal experience with Method and apparatus for extracting speech related facial featur, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Method and apparatus for extracting speech related facial featur will most certainly appreciate the feedback.
Profile ID: LFUS-PAI-O-1399951