Method of web crawling utilizing crawl numbers

Data processing: presentation processing of document – operator i – Presentation processing of document – Layout

Reexamination Certificate

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

C707S793000, C707S793000

Reexamination Certificate

active

06638314

ABSTRACT:

FIELD OF THE INVENTION
The present invention relates to the field of network information software and, in particular, to methods and systems for retrieving data from network sites.
BACKGROUND OF THE INVENTION
In recent years, there has been a tremendous proliferation of computers connected to a global network known as the Internet. A “client” computer connected to the Internet can download digital information from “server” computers connected to the Internet. Client application software executing on client computers typically accept commands from a user and obtain data and services by sending requests to server applications running on server computers connected to the Internet. A number of protocols are used to exchange commands and data between computers connected to the Internet. The protocols include the File Transfer Protocol (FTP), the Hyper Text Transfer Protocol (HTTP), the Simple Mail Transfer Protocol (SMTP), and the “Gopher” document protocol.
The HTTP protocol is used to access data on the World Wide Web, often referred to as “the Web.” The World Wide Web is an information service on the Internet providing documents and links between documents. The World Wide Web is made up of numerous Web sites located around the world that maintain and distribute electronic documents, A Web site may use one or more Web server computers that store and distribute documents in one of a number of formats including the Hyper Text Markup Language (HTML). An HTML document contains text and metadata or commands providing formatting information. HTML documents also include embedded “links” that reference other data or documents located on any Web server computers. The referenced documents may represent text, graphics, or video in respective formats.
A Web browser is a client application or operating system utility that communicates with server computers via FTP, HTTP, and Gopher protocols. Web browsers receive electronic documents from the network and present them to a user. Internet Explorer, available from Microsoft Corporation, of Redmond, Washington, is an example of a popular Web browser application.
An intranet is a local area network containing Web servers and client computers operating in a manner similar to the World Wide Web described above. Typically, all of the computers on an intranet are contained within a company or organization.
Web crawlers are computer programs that retrieve numerous electronic documents from one or more Web sites. A Web crawler processes the received data, preparing the data to be subsequently processed by other programs. For example, a Web crawler may use the retrieved data to create an index of documents available over the Internet or an intranet. A “search engine” can later use the index to locate electronic documents that satisfy a specified criteria.
A user that performs a document search provides search parameters to limit the number of documents retrieved. For example, a user may submit a search request that includes a list of one or more words, and the search engine locates electronic documents that contain a specified combination of the words. A user may repeat a search after a period of time. When a search is repeated, the user may prefer to avoid locating documents that have been located by prior searches.
It is desirable to have a mechanism by which a user can request a search engine to return only documents that have changed in some substantive way since that prior search. Preferably, such a mechanism will provide a Web crawler with a way to retrieve only documents that may have changed since a previous Web crawl and then to determine if an actual, substantive change has been made to the document. The mechanism would also preferably provide a way to mark the data retrieved from the document and stored in an index with an identifier that could be used in a search of the index to indicate when the Web crawler last found a substantive change to the document. The present invention is directed to providing such a mechanism.
SUMMARY OF THE INVENTION
In accordance with this invention, a system and computer based method of retrieving data from a computer network are provided. In an actual embodiment of the present invention, the method includes performing a Web crawl, by retrieving a set of electronic documents and subsequently retrieving additional electronic documents based on addresses specified within each electronic document. In a later Web crawl, electronic documents that have been modified subsequent to the previous Web crawl and electronic documents that were not retrieved during the previous Web crawl are retrieved. Electronic documents that were deleted since the previous Web crawl are detected. Each Web crawl is assigned a unique current crawl number. A crawl number modified is associated with and stored with the storage data from each electronic document retrieved during the Web crawl. The crawl number modified is set equal to the current crawl number when the document is first retrieved, or when it has previously been retrieved and has been found by the mechanism of the invention to have been modified in some substantive manner. In a subsequent search request, a crawl number can be retained as a search parameter and compared against a crawl number modified that is stored with the document data to determine if a document has been modified subsequent to the crawl number specified in the search.
In accordance with other aspects of this invention, each electronic document has a corresponding document address specification and provides information for locating the electronic document. During a Web crawl, document address specifications are used to retrieve copies of the corresponding electronic documents. Information from each electronic document retrieved during a Web crawl is stored in an index and associated with the corresponding document address specification and with a crawl number modified. If the retrieved document contains document address specifications to linked documents included in hyperlinks, these linked documents are also selectively retrieved during the Web crawl and processed in the manner described above.
In accordance with further aspects of this invention, performing a Web crawl includes assigning a unique current crawl number to the Web crawl, and determining whether a currently retrieved electronic document corresponding to each previously retrieved electronic document copy is substantively equivalent to the corresponding previously retrieved electronic document copy, in order to determine whether the electronic document has been modified since a previous crawl. If the current electronic document is not substantively equivalent to the previously retrieved electronic document copy, and therefore has been modified, the document's associated crawl number modified is set to the current crawl number and stored in the index with the data from the current electronic document.
In accordance with still other aspects of this invention, a secure hash function is used to determine a hash value corresponding to each retrieved electronic document copy. The hash value is stored in the index and used in subsequent Web crawls to determine whether the corresponding electronic document is modified. The current electronic document is retrieved and used to obtain a new hash value, which is compared with the previously determined hash value corresponding to the associated document address specification that is stored in a history map. If the hash values are equal, the current electronic document is considered to be substantively equivalent to the previously retrieved electronic document copy. If the hash values differ, the current electronic document is considered to be modified and the current crawl number is associated with the newly retrieved electronic document as the crawl number modified. The crawl number modified indicates the crawl number of the last crawl in which the data in the document was found to have changed. The hash value is stored with the associated data from the retrieved document and stored in the index. Preferabl

LandOfFree

Say what you really think

Search LandOfFree.com for the USA inventors and patents. Rate them and share your experience with other people.

Rating

Method of web crawling utilizing crawl numbers does not yet have a rating. At this time, there are no reviews or comments for this patent.

If you have personal experience with Method of web crawling utilizing crawl numbers, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Method of web crawling utilizing crawl numbers will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFUS-PAI-O-3125958

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.