| Along with the development of Internet, network information increases rapidly. In order to make the information service more efficient and true, we should reasonably get the information in Internet organized and classified. The thesis focuses on texts information processing in the Hypertext Markup Language and precedes the thorough research to texts classification from theory and application. The contributions of this dissertation are as follows:1. Build an experimental corpus.2. In the thesis, we investigate the function of HTML tags which is used to decorate the WebPages content based on prevenient theory, WebPages analysis and weighting tactics based on HTML tags are designed and realized.3. Technologies of WebPages categorization is analyzed, including: texts pretreatment, features weight, six feature evaluation functions about feature distillation and feature selection: Information Gain, Mutual Information, Expected Cross Entroph, X~2-statistic, the Weight of Evidence for .Text, Right half of IG Web corpus from Webdup are tested for evaluating functions of KNN and KNN-SVM categorization machines.4. Three categorization arithmetic of HTML texts, namely, Naive Bayes , KNN, and SVM are analyzed. Through combining KNN whit SVM, KNN-SVM is formed... |