Font Size: a A A

Automatic Classification Research On HTML Document And Implentation Of The Tool

Posted on:2007-05-27Degree:MasterType:Thesis
Country:ChinaCandidate:D M LiuFull Text:PDF
GTID:2178360185981912Subject:Computer application technology
Abstract/Summary:
Along with the development of Internet, network information increases rapidly. In order to make the information service more efficient and true, we should reasonably get the information in Internet organized and classified. The thesis focuses on texts information processing in the Hypertext Markup Language and precedes the thorough research to texts classification from theory and application. The contributions of this dissertation are as follows:1. Build an experimental corpus.2. In the thesis, we investigate the function of HTML tags which is used to decorate the WebPages content based on prevenient theory, WebPages analysis and weighting tactics based on HTML tags are designed and realized.3. Technologies of WebPages categorization is analyzed, including: texts pretreatment, features weight, six feature evaluation functions about feature distillation and feature selection: Information Gain, Mutual Information, Expected Cross Entroph, X~2-statistic, the Weight of Evidence for .Text, Right half of IG Web corpus from Webdup are tested for evaluating functions of KNN and KNN-SVM categorization machines.4. Three categorization arithmetic of HTML texts, namely, Naive Bayes , KNN, and SVM are analyzed. Through combining KNN whit SVM, KNN-SVM is formed...
Keywords/Search Tags:HTML Text Automatic Classifion, Vector Space Model, K-Nearest Neighber Classifier, Support Vector Machines, KNN-SVM
Related items