Font Size: a A A

Automatic Summarization Method Of Technical Literature Based On Domain Ontology

Posted on:2021-07-30Degree:MasterType:Thesis
Country:ChinaCandidate:P CaoFull Text:PDF
GTID:2518306107977229Subject:Computer Science and Technology
Abstract/Summary:
In recent years,with the rapid development of the Internet,new media methods bring mass data.How to get information from a large amount of text information more accurately and quickly is a hot topic in information science research.Summarization is the core content of a document,which can help people obtain the main information efficiently.This thesis first analyzes the current status of research on automatic summarization at home and abroad.Automatic summarization technology can be divided into mechanical and understanding summarization.There are two difficulties in current summarization techniques: 1)mechanical summarization methods cannot use the semantic information of documents fully.2)comprehension summarization cannot integrate the semantic information of long text.Because ontology is semantically complete and reasonable.Ontology has semantic completeness and deductibility.We propose an automatic summarization method based on domain ontology.However,the construction of domain ontology relies on domain experts,which greatly reduces the universality and extensibility of this method.Aiming at the problem of artificial dependency of domain ontology,a method of automatic construction of domain ontology is proposed.Through the above analysis,the main contents of this thesis are as follows:(1)Construct a set of domain ontology concepts.We obtain scientific and technological literature related to "automatic summarization" from the network knowledge base as the original training corpus of the domain ontology,and use TF-IDF algorithm to extract relevant keywords in the field of "automatic summarization".At the same time,using Word2 Vector model to get the word vector set.(2)Get the relationship of domain ontology.Screen and revise according to the reserved concepts and the existing knowledge graphs to obtain the relationship between hierarchical concepts of domain ontology.Combining FP-Growth algorithm to mine the relationship between concepts in the domain text in recent years to increase the semantic coverage and timeliness of the domain ontology.(3)Construct the inference rule library and invoke the Jena inference machine to map the document sentences to RDF triples for semantic inference.Improve the semantic accuracy of text mapping results.(4)In order to verify the distribution of summarize sentences extracted by this method from various documents and paragraphs,an evaluation index for summary coverage and uniformity was proposed.The evaluation criteria of the summarization is ROUGE: experiments show that the quality of the document summarization based on domain ontology is higher than that of the traditional method.The ROUGE value of multi-document summarizations obtained by this method is 15% higher than that of single-document summarizations.The result of uniformity calculation shows that the summarization generated by using domain ontology is more comprehensive and uniform than the traditional method.With the help of the document summarization derived from the domain ontology,the experiment can fully mine the semantic information between long documents and multiple documents and accurately express the main idea of the text.
Keywords/Search Tags:Domain Ontology, Automatic Summarization, Association Rule Mining, RDF, Knowledge Graph
Related items