Methods and apparatus for identifying unknown speakers using...

Data processing: speech signal processing – linguistics – language – Speech signal processing – Recognition

Reexamination Certificate

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

C704S248000, C704S250000

Reexamination Certificate

active

06748356

ABSTRACT:

FIELD OF THE INVENTION
The present invention relates generally to audio information classification systems and, more particularly, to methods and apparatus for identifying speakers in an audio file.
BACKGROUND OF THE INVENTION
Many organizations, such as broadcast news organizations and information retrieval services, must process large amounts of audio information, for storage and retrieval purposes. Frequently, the audio information must be classified by subject or speaker name, or both. In order to classify audio information by subject, a speech recognition system initially transcribes the audio information into text for automated classification or indexing. Thereafter, the index can be used to perform query-document matching to return relevant documents to the user.
Thus, the process of classifying audio information by subject has essentially become fully automated. The process of classifying audio information by speaker, however, often remains a labor intensive task, especially for real-time applications, such as broadcast news. While a number of computationally-intensive off-line techniques have been proposed for automatically identifying a speaker from an audio source using speaker enrollment information, the speaker classification process is most often performed by a human operator who identifies each speaker change, and provides a corresponding speaker identification.
A number of techniques have been proposed or suggested for identifying speakers in an audio stream, including U.S. patent application Ser. No. 09/434,604, filed Nov. 5, 1999, U.S. patent application Ser. No. 09/345,237, filed Jun. 30, 1999, and U.S. patent application Ser. No. 09/288,724, filed Apr. 9, 1999, each assigned to the assignee of the present invention. U.S. patent application Ser. No. 09/345,237, for example, discloses a method and apparatus for automatically transcribing audio information from an audio source while concurrently identifying speakers in real-time, using an existing enrolled speaker database. U.S. patent application Ser. No. 09/345,237, however, can only identify the set of the speakers in the enrolled speaker database. In addition, U.S. patent application Ser. No. 09/345,237 does not allow new speakers to be added to the enrolled speaker database while audio information is being processed in real-time. A need therefore exists for a method and apparatus that automatically identifies unknown speakers in real-time or in an off-line manner. A further need exists for a method and apparatus that automatically identifies unknown speakers using a hierarchical speaker model tree.
SUMMARY OF THE INVENTION
Generally, a method and apparatus are disclosed for identifying speakers participating in an audio-video source, whether or not such speakers have been previously registered or enrolled. A speaker segmentation system separates the speakers and identifies all possible frames where there is a segment boundary between non-homogeneous speech portions. The hierarchical speaker tree clustering system clusters homogeneous segments (generally corresponding to the same speaker), and assigns a cluster identifier to each detected segment, whether or not the actual name of the speaker is known. Thus, segments corresponding to the same speaker should have the same cluster identifier.
According to one aspect of the invention, the disclosed speaker identification system uses a hierarchical enrolled speaker database that includes one or more background models for unenrolled speakers to assign a speaker to each identified segment. Once the speech segments are identified by the segmentation system, the disclosed unknown speaker identification system compares the segment utterances to the enrolled speaker database using a hierarchical approach and finds the “closest” speaker, if any, to assign a speaker label to each identified segment.
A speech segment having an unknown speaker is initially assigned a general speaker label from a set of background models for speaker identification, such as “unenrolled male” or “unenrolled female.” The “unenrolled” segment is assigned a cluster identifier and is positioned in the hierarchical tree. Thus, the hierarchical speaker tree clustering system assigns a unique cluster identifier corresponding to a given node, for each speaker to further differentiate the general speaker labels.


REFERENCES:
patent: 5659662 (1997-08-01), Wilcox et al.
patent: 6345252 (2002-02-01), Beigi et al.
patent: 6421645 (2002-07-01), Beigi et al.
patent: 6424946 (2002-07-01), Tritschler et al.
patent: 6684186 (2004-01-01), Beigi et al.
S. Dharanipragada et al., “Experimental Results in Audio Indexing,” Proc. ARPA SLT Workshop, (Feb. 1996).
L. Polymenakos et al., “Transcription of Broadcast News—Some Recent Improvements to IBM's LVCSR System,” Proc. ARPA SLT Workshop, (Feb. 1996).
R. Bakis, “Transcription of Broadcast News Shows with the IBM Large Vocabulary Speech Recognition System,” Proc. ICASSP98, Seattle, WA (1998).
H. Beigi et al., “A Distance Measure Between Collections of Distributions and its Application to Speaker Recognition” Proc. ICASSP98, Seattle, WA (1998).
S. Chen, “Speaker, Environment and Channel Change Detection and Clustering via the Bayesian Information Criterion,” Proceedings of the Speech Recognition Workshop (1998).
S. Chen et al., “Clustering via the Bayesian Information Criterion with Applications in Speech Recognition,” Proc. ICASSP98, Seattle, WA (1998).
S. Chen et al., “IBM's LVCSR System for Transcription of Broadcast News Used in the 1997 Hub4 English Evaluation,” Proceedings of the Speech Recognition Workshop (1998).
S. Dharanipragada et al., “A Fast Vocabulary Independent Algorithm for Spotting Words in Speech,” Proc. ICASSP98, Seattle, WA (1998).
J. Navratil et al., “An Efficient Phonotactic-Acoustic system for Language Identification,” Proc. ICASSP98, Seattle, WA (1998).
G. N. Ramaswamy et al., “Compression of Acoustic Features for Speech Recognition in Network Environments,” Proc. ICASSP98, Seattle, WA (1998).
S. Chen et al., “Recent Improvements to IBM's Speech Recognition System for Automatic Transcription of Broadcast News,” Proceedings of the Speech Recognition Workshop (1999).
S. Dharanipragada et al., “Story Segmentation and Topic Detection in the Broadcast News Domain,” Proceedings of the Speech Recognition Workshop (1999).
C. Neti et al., “Audio-Visual Speaker Recognition for Video Broadcast News,” Proceedings of the Speech Recognition Workshop (1999).

LandOfFree

Say what you really think

Search LandOfFree.com for the USA inventors and patents. Rate them and share your experience with other people.

Rating

Methods and apparatus for identifying unknown speakers using... does not yet have a rating. At this time, there are no reviews or comments for this patent.

If you have personal experience with Methods and apparatus for identifying unknown speakers using..., we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Methods and apparatus for identifying unknown speakers using... will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFUS-PAI-O-3362755

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.