Method, apparatus, and computer program product for...

Data processing: speech signal processing – linguistics – language – Linguistics – Natural language

Reexamination Certificate

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

C707S793000

Reexamination Certificate

active

06338034

ABSTRACT:

BACKGROUND OF THE INVENTION
1. Field of the Invention
The present invention relates to a method, an apparatus, and a computer program product for making a summary of a document on the basis of commonly owned information and relation between ideas found in the sentences in the document.
2. Description of the Prior Art
In a computer-aided automatic summarizing system, important sentences are selected from the sentences in a document according to an important-sentence selecting rule, and then, a summary is made from the important sentences according to a summarizing rule.
Examples of summarizing techniques using an important-sentence selecting rule are disclosed in Japanese Patent Laid-Open Publication Nos. 2-93866 and 2-112069. In these examples, keywords are selected among words in a document based on the frequency distribution of occurrence of the same word, database of keywords, and user's decision, and then, important sentences containing the keywords are selected. Japanese Patent Laid-Open Publication No. 2-215049 discloses a technique to select important parts from a document by analyzing context vectors. Moreover, Japanese Patent Laid-Open Publication No. 3-191475 discloses a technique to select important sentences from a document by applying a rule reflecting the feature of each paragraph to sentences therein.
The technique of automatically making a summary of a document by gathering sentences containing keywords inevitably requires the use of database of keywords in making a decision as to whether or not each sentence contains keywords, and therefore, the resultant summary is likely to be affected by the contents of the database, and the contents of the summary may be biased or lack flexibility.
The technique of automatically making a summary of a document on the basis of the structure of the document such as the context thereof or the feature of the paragraphs therein is suitable for selecting the structure of a long document and catching the transition of the subject of a long document, but is not suitable for creating a compact summary of high-accuracy.
The technique of making a summary from a collection of important sentences in a document requires an effective method for collecting information common to the important sentences and generating a compact expression reflecting the common information. However, such method has not been proposed as yet.
SUMMARY OF THE INVENTION
An object of the present invention is to provide a method, an apparatus, and a computer program product for generating from a document an accurate and compact summary on which the point of view of a user is reflected.
According to the present invention, there is provided method of summarizing a document which comprises the steps of: extracting sentence-constituting-elements from the document; tabularizing the sentence-constituting-elements corresponding to categories and sentences in the document; extracting commonly-held-information which is common to the sentence-constituting-elements in the same category from the sentence-constituting-elements; looking up common expression information which is common to plural pieces of the commonly-held-information in a thesaurus in which the commonly-held-information and the common expression information are connected by a hierarchical tree; and composing a summary based on the commonly-held-information and the common expression information.
Sentence-constituting-elements which come under predetermined categories are extracted from a document. The predetermined categories includes “When”, “Where”, “Who”, “What”, “Why”, “How” (5W1H) and “Done”. The extracted sentence-constituting-elements are categorized based on the categories. The categorized sentence-constituting-elements are referred to as categorized information. The categorized information which is subjected to generation of the framework of a sentence in a summary is referred to as commonly-held-information. The common expression information which is common to commonly-held-information is extracted consulting a thesaurus. The categorized information which is subjected to generation of the qualifying part of a sentence in the summary is referred to as commonly-occurring-information. A sentence in the summary is generated from pieces of commonly-held-information, pieces of common expression information and optionally from pieces of commonly-occurring-information.
These and other objects, features and advantages of the present invention will become more apparent in light of the following detailed description of the best mode embodiments thereof, as illustrated in the accompanying drawings.


REFERENCES:
patent: 4965763 (1990-10-01), Zamora
patent: 5077668 (1991-12-01), Doi
patent: 5297027 (1994-03-01), Morimoto et al.
patent: 5638543 (1997-06-01), Pedersen et al.
patent: 5689716 (1997-11-01), Chen
patent: 5778397 (1998-07-01), Kupiec et al.
patent: 5838323 (1998-11-01), Rose et al.
patent: 5848191 (1998-12-01), Chen
patent: 5857184 (1999-01-01), Lynch
patent: 5873087 (1999-02-01), Brosda et al.
patent: 5918240 (1999-07-01), Kupiec et al.
patent: 5924108 (1999-07-01), Fein et al.
patent: 5963965 (1999-10-01), Vogel
patent: 1-290076 (1989-11-01), None
patent: 2-93866 (1990-04-01), None
patent: 2-112069 (1990-04-01), None
patent: 3-191475 (1991-08-01), None
patent: 4-74259 (1992-03-01), None
patent: 4-90055 (1992-03-01), None
patent: 5-233729 (1993-09-01), None
patent: 6-215049 (1994-08-01), None
patent: 6-259423 (1994-09-01), None
patent: 7-175808 (1995-07-01), None
Joachim, Jae-Hak et al., “Another Investigation of Automatic Text Summarization: A Reader Oriented Approach”, Intelligent Information Systems 1994, Proceeding of the 1994 Second Australian and New Zealand Conference, 1994, pp. 472-476.*
Doi Nishimura “Japanese Language Text Summary” NEC Technical Report, vol. 47 No. 8, 1994, pp. 48-52 (Sep. 16, 1994).
Ando, Doi, Muraki “Information Extraction from Newspaper Articles and Providing a Multi-Language Index”, Information Processing Society of Japan 48th(first half of 1994) National Conference Proceedings (3), p. 105-106 (Mar. 23, 1994).

LandOfFree

Say what you really think

Search LandOfFree.com for the USA inventors and patents. Rate them and share your experience with other people.

Rating

Method, apparatus, and computer program product for... does not yet have a rating. At this time, there are no reviews or comments for this patent.

If you have personal experience with Method, apparatus, and computer program product for..., we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Method, apparatus, and computer program product for... will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFUS-PAI-O-2822980

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.