Home /Research /A gleaning subsystem for CINDI
OTHER

A gleaning subsystem for CINDI

Tong Zhang

Year
2004
Citations
2
Access
Open access

Abstract

Internet search engines typically use Internet crawlers, or robots, for the purpose of constructing and maintaining a searchable index of resources on the Web. Topic-specific robots will become popular in the next generation. They gather information on the Internet in specific domains by means of information filtering technology. The CINDI Robot System is such an application in academic domain. This research is concerned with a structure-based gleaning subsystem for CINDI. The system separates theses, technical reports, academic papers, and FAQs as resources while e-mails, letters, resumes, graphics, and discussion groups are considered as chaff. This system makes decisions based on weight, which is carefully assigned to each resource by matching its structure with predefined Document Type Definitions (DTDs). The DTDs for the typical structure for the specific document types are built based on some predefined profiles. The system also features conversion subsystem in Windows environment to unify document formats for CINDI. (Abstract shortened by UMI.)

Keywords

The InternetComputer scienceInformation retrievalResource (disambiguation)World Wide WebGraphicsDomain (mathematical analysis)Computer graphics (images)

Related papers

Browse all OTHER papers