Font Size: a A A

Transportation Facility Information Extraction Based On Web Mining

Posted on:2023-02-03Degree:MasterType:Thesis
Country:ChinaCandidate:L B ZhengFull Text:PDF
GTID:2530307073485394Subject:Surveying the science and technology
Abstract/Summary:
With the rapid development of computer information technology and network technology,the trend of digitalization and intelligence of modern transportation is unstoppable.For the traditional traffic facility information,due to the incomplete information of early remote sensing image samples and other attributes,there are problems of large sample size and complicated manual annotation,which have gradually failed to meet the development needs of intelligent transportation as well as to cope with the possible technical challenges in the future.Therefore,there is an increasing need to automatically establish geosyntactic object-centered multi-source data association from a large amount of web-source information,and then realize automatic annotation of geographic objects.This thesis focuses on a Web-oriented traffic facility attribute information extraction method based on the huge amount of relevant traffic facility attribute information already existing in the Internet,which can quickly,comprehensively and reliably automatically mine,extract and store traffic facility information from the Web and combine it with demand for rational application.The following research work is carried out in this thesis:(1)Acquisition and screening of related traffic facilities web pages: A method for automatically acquiring web pages related to the target traffic facilities is proposed;screening criteria based on web page content topics are proposed,and an in-network topic content discrimination method based on key sentences of the main text of the web page is also proposed.(2)Extraction of traffic facility information for Web pages: the traffic facilityrelated attribute information is specified,and the extraction range of traffic facility information content is determined;a set of Web page information extraction methods for multi-source information structure and the processing process after information extraction are designed by combining the traffic information attribute characteristics and the types and existence forms of information in Web pages;the reliability of the extracted information is verified through reliability evaluation The validity of the extracted information and the scientificity of the extraction method were verified by reliability evaluation.(3)Organization,storage and management of extracted information: The data storage structure of non-relational traffic facility information is designed by making full use of the flexibility and scalability of Mongo DB document database and using the collection of single traffic facility information as the basic storage unit for the extracted information.The possible needs of data in storage use are analyzed in detail,and the data reading and writing functions are designed considering the actual management needs of users.(4)Based on the above technical solutions,an online Web page traffic information extraction system is developed around the analysis of requirements,structure and function design,which can automate the functions of traffic facility object creation,traffic facility information extraction and editing,and traffic facility information export,etc.The feasibility and scientificity of the system are proved by the practical application results.The research results of this thesis can be effectively applied to the work of traffic facility attribute information acquisition,and its related technology can serve the construction of public transportation,play the role of information complementation and information and quality assessment traceability.
Keywords/Search Tags:traffic facilities, Web mining, information extraction, text classification, Data Management
Related items