Font Size: a A A

Research And Application Of Semantic-Based Information Extraction

Posted on:2012-11-14Degree:MasterType:Thesis
Country:ChinaCandidate:H E ZhangFull Text:PDF
GTID:2218330362454337Subject:Computer software and theory
Abstract/Summary:
World Wide Web is one of the largest knowledge base for public information, which contains a great number of information. How to extract information which meets the customer needs is an important point in the engineering field. At present the Web information extraction is based on keyword or HTML, it calculates sorts and indexes. These methods are based on pattern matching syntax, it cannot change rules automatically when the keyword or HTML changes. The search engines cannot understand the semantics of the search semantic, and the semantics of this information field, it is trying to help the user find the information, but also to determine which the best information is. Search engines to search the Web information through semantic is very difficult, because that most of the information on the Web is completely user-readable and human understandable. The key to solve problems is that designing a correct and effective semantic matching method.Based on the existing information extraction method, this paper has mainly completed the following tasks:①Compared and analyzed the method of HTML document to XML document conversion, add semantic computation to conversion, it improved the conversion of the list-based method, and the accuracy of the document conversion.②According to the futures field leads to semantic ambiguity, it designed a futures domain ontology using ontology learning methods and Protégémodeling tools, it solved the heterogeneity.③Based on the traditional semantic similarity algorithm, improved semantic similarity computation, proposed the hierarchy-based algorithm, this improved the reasonable of similarity computation,④Based on the method above, proposed a semantic-based Web information extraction method, and verify the correctness of their approach.⑤Based on proposed information extraction, designed and implemented semantic-based Web information extraction system, used it in the analysis of the future goods, verified the feasibility and availability.
Keywords/Search Tags:HTML, Ontology, Semantic, XML, Information Extraction
Related items