Automatic rating and filtering of data files for...

Electrical computers and digital processing systems: multicomput – Distributed data processing – Client/server

Reexamination Certificate

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

C709S228000, C709S229000, C707S793000, C707S793000, C707S793000, C725S028000, C386S349000

Reexamination Certificate

active

06493744

ABSTRACT:

FIELD OF THE INVENTION
This invention relates generally to methods for rating data for objectionable content. More particularly, it relates to methods for automatically rating and filtering objectionable data on Web pages.
BACKGROUND ART
The astronomical growth of the World Wide Web in the last decade has put a wide variety of information at the fingertips of anyone with access to a computer connected to the internet. In particular, parents and teachers have found the internet to be a rich educational tool for children, allowing them to conduct research that would in the past have either been impossible or taken far too long to be feasible. In addition to valuable information, however, children also have access to offensive or inappropriate information, including violence, pornography, and hate-motivated speech. Because the World Wide Web is inherently a forum for unrestricted content from any source, censoring material that some find objectionable is an unacceptable solution.
Voluntary user-based solutions have been developed for implementation with a Web browser on a client computer. The browser determines whether or not to display a document by applying a set of user-specified criteria. For example, the browser may have access to a list of excluded sites or included sites, provided by a commercial service or a parent or educator. Users can also choose to receive documents only through a Web proxy server, which compares the requested document with an exclusion or inclusion list before sending it to the client computer. Because new content is continually being added to the World Wide Web, however, it is virtually impossible to maintain a current list of inappropriate sites. Limiting the user to a list of included sites might be appropriate for corporate environments, but not for educational ones in which the internet is used for research purposes.
The Recreational Software Advisory Council (RSAC) has developed an objective content rating labeling system for Web sites, called RSAC on the Internet (RSACi). The system produces ratings tags that are compliant with the Platform for Internet Content Selection (PICS) tag system already in place, and that can easily be incorporated into existing HTML documents. The RSACi labels rate content on a scale of zero to four in four categories: violence, nudity, sex, and language. Current Web browsers are designed to read the RSACi tags and determine whether or not to display the document based on content levels the user sets for each of the four categories. The user can also set the browser not to display pages without a rating.
While a good beginning, there are three significant limitations to the RSACi rating system. First, it is a voluntary system and is effective only if widely implemented. There is somewhat of an incentive for the site creator to assign a rating, even if a zero rating, because some users choose not to display sites without a rating. If the site's creator does not include a rating, it can be generated by an outside source. However, the rate at which content is being added to the Web makes it virtually impossible for a third party to rate every new Web site manually.
Second, while the RSACi rating aims to be objective, it is subject to some amount of discretion of the person doing the rating. At its Web site (http://www.rsac.org), RSAC provides a detailed questionnaire for providing the rating, but the user can easily override or adjust the results.
Finally, there is currently no way to rate dynamically created documents. For example, search engines receive a user query, find applicable documents, and create a search result page listing a number of the located documents. The search result page typically includes a title and short abstract or extract, along with the URL, for each retrieved document. The result page itself might have objectionable content, and currently the only way to address this problem is for browsers not to display search result pages at all. Without search engines, though, internet research is significantly limited.
A further problem with all of the above solutions, as well as with word-screening or phrasescreening systems, is that they either allow or deny access to Web pages. Even if only a small portion of the document is objectionable, the user is prohibited from seeing the entire document. This is especially significant in search result pages, in which one offensive site prevents display of all of other unrelated sites.
The situation becomes even more complex when Web pages include non-text data, for example, audio or images. Surrounding text does not always indicate the content of the embedded file, allowing offensive audio or image material to slip through the ratings system. Occasionally, people deliberately mislabel offensive audio or image files in order to mislead monitoring services.
There is a need, therefore, for an automatic rating method for all material available on the World Wide Web, including dynamically created material, that allows greater viewer control over what material is displayed or blocked.
OBJECTS AND ADVANTAGES
Accordingly, it is a primary object of the present invention to provide a method for automatically rating a data file, for example, a Web page, for objectionable content.
It is an additional object of the invention to provide an objective rating method that requires no subjective human input after the system is initially devised.
It is a further object of the present invention to provide a method for automatically rating dynamically created documents as they are being created.
It is a yet another object of the present invention to provide a rating and filtering method that blocks objectionable content of a file while allowing access to remaining inoffensive portions of the file.
It is an additional object of the present invention to provide a method that can be used with any type of data file, including text, audio, and image.
It is a further object to provide a method for rating and filtering data files that can be implemented on a client, server, or proxy server, and can therefore be easily incorporated into existing system architectures.
Finally, it is an object of the present invention to provide an automatic rating method that works with existing manual rating methods and requires minimal system changes.
SUMMARY
These objects and advantages are attained by a computer-implemented method for rating a raw data file for objectionable content. The method occurs in a distributed computer system and comprises the steps of preprocessing the raw data file to create semantic units representative of the semantic content of the raw data file, comparing the semantic units with a rating repository comprising semantic entries and corresponding ratings, assigning content rating vectors to the semantic units, and creating a modified data file incorporating rating information derived from the content rating vectors. After the modified data file is created, either all, some, or none of the file will be displayed by a browser to a user at a client computer.
The method works with any type of data file that can be converted to semantic units. Embodiments of the preprocessing step vary with the type of raw data file to be rated. In one embodiment, a text-only HTML document is stripped of its tags and is then parsed into semantic units, for example, words or phrases. In an alternate embodiment, the data file is an audio file, and text data is created from the audio file using standard voice recognition software. The system also creates an audio-to-text correlation between a location in the created text data and a corresponding location in the audio file. The text file is then parsed into semantic units. In a further embodiment, image processing software is used to identify semantic units within an image file. The semantic units of an image file are discrete objects in regions within the image file.
The rating repository used depends on the type of file and related semantic units. For text files, the repository contains entries of words or phrases with corresponding content ra

LandOfFree

Say what you really think

Search LandOfFree.com for the USA inventors and patents. Rate them and share your experience with other people.

Rating

Automatic rating and filtering of data files for... does not yet have a rating. At this time, there are no reviews or comments for this patent.

If you have personal experience with Automatic rating and filtering of data files for..., we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Automatic rating and filtering of data files for... will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFUS-PAI-O-2920843

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.