System and method of mu-law or A-law compression of bark...

Data processing: speech signal processing – linguistics – language – Speech signal processing – Recognition

Reexamination Certificate

Rate now

[ 0.00 ] – not rated yet Voters 0 Comments 0

Details System and method of mu-law or A-law compression of bark... System and method of mu-law or A-law compression of bark...

Filed on: 2000-10-31
Date issued: 2004-02-17
Examiner: Dorvil, Richemond (Department: 2654)
Patent Class: Data processing: speech signal processing, linguistics, language
SubClass: Speech signal processing
SubSubClass: Recognition

Other Related Categories: C704S200100
Type: Reexamination Certificate
Status: active
Patent number: 06694294
Description: ABSTRACT:

BACKGROUND
I. Field
The present invention pertains generally to the field of communications and more specifically to a system and method for improving voice recognition in noisy environments and frequency mismatch conditions.
II. Background
Voice recognition (VR) represents one of the most important techniques to endow a machine with simulated intelligence to recognize user or user-voiced commands and to facilitate human interface with the machine. VR also represents a key technique for human speech understanding. Systems that employ techniques to recover a linguistic message from an acoustic speech signal are called voice recognizers. The term “voice recognizer” is used herein to mean generally any spoken-user-interface-enabled device.
The use of VR (also commonly referred to as speech recognition) is becoming increasingly important for safety reasons. For example, VR may be used to replace the manual task of pushing buttons on a wireless telephone keypad. This is especially important when a user is initiating a telephone call while driving a car. When using a phone without VR, the driver must remove one hand from the steering wheel and look at the phone keypad while pushing the buttons to dial the call. These acts increase the likelihood of a car accident. A speech-enabled phone (i.e., a phone designed for speech recognition) would allow the driver to place telephone calls while continuously watching the road. In addition, a hands-free car-kit system would permit the driver to maintain both hands on the steering wheel during call initiation.
Speech recognition devices are classified as either speaker-dependent (SD) or speaker-independent (SI) devices. Speaker-dependent devices, which are more common, are trained to recognize commands from particular users. In contrast, speaker-independent devices are capable of accepting voice commands from any user. To increase the performance of a given VR system, whether speaker-dependent or speaker-independent, training is required to equip the system with valid parameters. In other words, the system needs to learn before it can function optimally.
An exemplary vocabulary for a hands-free car kit might include the digits on the keypad; the keywords “call,” “send,” “dial,” “cancel,” “clear,” “add,” “delete,” “history,” “program,” “yes,” and “no”; and the names of a predefined number of commonly called coworkers, friends, or family members. Once training is complete, the user can initiate calls by speaking the trained keywords, which the VR device recognizes by comparing the spoken utterances with the previously trained utterances (stored as templates) and taking the best match. For example, if the name “John” were one of the trained names, the user could initiate a call to John by saying the phrase “Call John.” The VR system would recognize the words “Call” and “John,” and would dial the number that the user had previously entered as John's telephone number. Garbage templates are used to represent all words not in the vocabulary.
Combining multiple engines provides enhanced accuracy and uses a greater amount of information in the input speech signal. A system and method for combining VR engines is described in U.S. patent application Ser. No. 09/618,177 (hereinafter '177 application) entitled “Combined Engine System and Method for Voice Recognition”, filed Jul. 18, 2000, and U.S. patent application Ser. No. 09/657,760 (hereinafter '760 application) entitled “System and Method for Automatic Voice Recognition Using Mapping,” filed Sep. 8, 2000, which are assigned to the assignee of the present invention and fully incorporated herein by reference.
Although a VR system that combines VR engines is more accurate than a a VR system that uses a singular VR engine, each VR engine of the combined VR system may include inaccuracies because of a noisy environment. An input speech signal may not be recognized because of background noise. Background noise may result in no match between an input speech signal and a template from the VR system's vocabulary or may cause a mismatch between an input speech signal and a template from the VR system's vocabulary. When there is no match between the input speech signal and a template, the input speech signal is rejected. A mismatch results when a template that does not correspond to the input speech signal is chosen by the VR system. The mismatch condition is also known as substitution because an incorrect template is substituted for a correct template.
An embodiment that improves VR accuracy in the case of background noise is desired. An example of background noise that can cause a rejection or a mismatch is when a cell phone is used for voice dialing while driving and the input speech signal received at the microphone is corrupted by additive road noise. The additive road noise may degrade voice recognition and accuracy and cause a rejection or a mismatch.
Another example of noise that can cause a rejection or a mismatch is when the speech signal received at a microphone placed on the visor or a headset is subjected to convolutional distortion. Noise caused by convolutional distortion is known as convolutional noise and frequency mismatch. Convolutional distortion is dependent on many factors, such as distance between the mouth and microphone, frequency response of the microphone, acoustic properties of the interior of the automobile, etc. Such conditions may degrade voice recognition accuracy.
Traditionally, prior VR systems have included a RASTA filter to filter convolutional noise. However, background noise was not filtered by the RASTA filter. Thus, there is a need for a technique to filter both convolutional noise and background noise. Such a technique would improve the accuracy of a VR system.
SUMMARY
The described embodiments are directed to a system and method for improving the frontend of a voice recognition system. In one aspect, a system and method for voice recognition includes mu-law compression of bark amplitudes. In another aspect, a system and method for voice recognition includes A-law compression of bark amplitudes. Both mu-law and A-law compression of bark amplitudes reduce the effect of noisy environments, thereby improving the overall accuracy of a voice recognition system.
In another aspect, a system and method for voice recognition includes mu-law compression of bark amplitudes and mu-law expansion of RelAtive SpecTrAl (RASTA) filter outputs. In yet another aspect, a system and method for voice recognition includes A-law compression of bark amplitudes and A-law expansion of RASTA filter outputs. When mu-law compression and mu-law expansion, or A-law compression and A-law expansion, are used, a matching engine such as a Dynamic Time Warping (DTW) engine is better able to handle channel mismatch conditions.

REFERENCES:
patent: 5450522 (1995-09-01), Hermansky et al.
patent: 5475792 (1995-12-01), Stanford et al.
patent: 5537647 (1996-07-01), Hermansky et al.
patent: 5615296 (1997-03-01), Stanford et al.
patent: 5754978 (1998-05-01), Perez-Mendez
patent: 6044340 (2000-03-01), Van Hamme
patent: 6092039 (2000-07-01), Zingher
patent: 19710953 (1997-07-01), None
Openshaw, J. P., Z. P. Sun, and J.S. Mason, “A Comparison of Composites Features Under Degraded Speech in Speaker Recognition,” IEEE Int. Conf. Acoust., Speech, and Sig. Proc., 1993 ICASSP-93, Apr. 27-30, 1993, vol. 2, pp. 371-374.*
U.S. application Ser. No. 09/657,760, entitled “System and Method For Automatic Voice Recognition Using Mapping,” filed Sep. 8, 2000. Ning Bi, et al., Qualcomm Inc., San Diego, California (USA).

Affiliated with

Garudadri Harinath

Inventor

[ 0.00 ] – not rated yet Voters 0 Comments 0

Also associated with

Brown Charles D.

Attorney

[ 0.00 ] – not rated yet Voters 0 Comments 0

Dorvil Richemond

Examiner

[ 0.00 ] – not rated yet Voters 0 Comments 0

Pappas George C.

Attorney

[ 0.00 ] – not rated yet Voters 0 Comments 0

Qualcomm Incorporated

Corporate Assignee

[ 0.00 ] – not rated yet Voters 0 Comments 0

Storm Donald L.

Examiner

[ 0.00 ] – not rated yet Voters 0 Comments 0

LandOfFree

Say what you really think

Search LandOfFree.com for the USA inventors and patents. Rate them and share your experience with other people.

Rating

System and method of mu-law or A-law compression of bark... does not yet have a rating. At this time, there are no reviews or comments for this patent.
If you have personal experience with System and method of mu-law or A-law compression of bark..., we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and System and method of mu-law or A-law compression of bark... will most certainly appreciate the feedback.

Rate now

Comments { 0 }

Profile ID: LFUS-PAI-O-3283156

All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.

Canada

Charities
Companies
MP Candidates
Patents
Employee Salary Disclosure

World

Places of the World
Scientific Papers

United States

Banks
Companies
Counties
Patents
Employee Salary Disclosure