Font Size: a A A

Research On Multi-source Document Metadata Fusion In Resource Discovery System

Posted on:2022-05-28Degree:MasterType:Thesis
Country:ChinaCandidate:X LiFull Text:PDF
GTID:2558306347990679Subject:Library and file management
Abstract/Summary:
With the rapid development of modern science and technology,the explosive growth and dissemination of scientific documents,the traditional form of document dissemination with paper text as the main carrier has seriously hindered the development of science and technology and the dissemination of scientific knowledge.The concept of digital resources and "digital library" is put forward.The collection form and service mode of traditional physical libraries are difficult to meet the needs of real development.However,the resource discovery system uses simple and efficient retrieval methods and high-capacity,low-cost academic resource storage methods.Adapted to the needs of academic resource development,injected new vitality into the development of traditional libraries,provided a new business model for various document resource service providers,and also brought users a brand-new document resource retrieval experience,becoming knowledge An important carrier for storage and dissemination.In the resource discovery system,metadata plays a very important role.It is the link that connects users,systems,and resources.On the one hand,metadata can help users quickly find target resources.On the other hand,service providers are performing resource discovery.During system construction and maintenance,metadata is the basis for the design of resource discovery service functions such as access point setting,faceted search,resource relevance judgment,and resource recommendation.It also plays a key role in the management and utilization of system database resources.However,the reality is that there are various problems with metadata in the database of each resource discovery system.In the process of organization,metadata providers have problems with tagging errors,permissions issues leading to incomplete metadata collected,and human description errors,etc.,causing various resources to find that the system metadata information varies in thickness and quality.At the same time,due to the different system function construction goals,the metadata heterogeneity of each resource discovery system is serious,and the interoperability between the resource discovery systems is poor.The above problems have brought great obstacles to the utilization and maintenance of metadata,and have increased the difficulty of metadata construction and maintenance for various resource discovery systems,and hindered the healthy development of resource discovery systems.In response to the above problems,many scholars have proposed various metadata quality control methods,but most of them are designed based on the process.Relying on single-source metadata to adopt process control methods cannot fundamentally solve the problems of metadata.Therefore,this article From the perspective of multi-source data fusion,a multi-source document metadata fusion model in the resource discovery system is proposed.On the one hand,it provides a new method for multi-source metadata fusion and enriches the relevant theories of metadata fusion.On the one hand,it provides guidance and reference significance for the improvement of metadata quality for various resource discovery systems,which can improve the level of metadata management and utilization,and lay a data foundation for a document discovery service with a good user experience.Based on the resource discovery system research and literature metadata sampling analysis,this paper summarizes two metadata fusion barriers:serious metadata heterogeneity and unsatisfactory metadata quality.Based on the fusion barriers,this paper puts forward the multi-source metadata in the resource discovery system.The fusion model divides the model into four links:metadata perspective,metadata preprocessing,metadata judgment,and metadata content fusion.Among them,the metadata judgment and the metadata fusion strategy model are the two most important links.The choice of the two strategies directly affects the final metadata fusion effect.In the selection of metadata judgment strategy,this paper proposes a metadata judgment strategy based on "title"similarity as the main framework.The consistency check of a few metadata items is used to realize the judgment of metadata of each source document.Metadata from each source can be merged after passing the metadata judgment strategy.When selecting the metadata content fusion strategy,according to the different situations that metadata fusion may face,this paper proposes 4 different content fusion strategies,respectively They are:deduplication fusion strategy,complementary fusion strategy,rule-based fusion strategy and weight-based fusion strategy.Finally,some journal papers,dissertations,conference papers metadata obtained from CNKI and Wanfang databases,and some library metadata obtained from the library of Central China Normal University and Hubei University of Technology were used to experiment with the model.The experimental results show that the accuracy and recall rate of the four types of document resource classification strategies are high,and the accuracy of the metadata content fusion strategy is also high.After content fusion,the quality of each metadata item is obtained.A certain percentage increase.From the experimental results,the metadata judgment strategy and metadata content fusion strategy designed in this paper are more effective,which verifies the feasibility of the multi-source document metadata fusion model proposed in this paper and its adaptability to multiple document types.
Keywords/Search Tags:Resource discovery system, Judging metadata duplication, Metadata fusion, Document type, Information entropy
Related items