Font Size: a A A

Research On Key Technologies Of Documents Reconstruction For Complex Image Documents

Posted on:2024-07-12Degree:DoctorType:Dissertation
Country:ChinaCandidate:G B WuFull Text:PDF
GTID:1528306944966749Subject:Information and Communication Engineering
Abstract/Summary:
The ubiquity of electronic documents has transformed the way we store,edit,and share information.However,in reality,a large number of paper documents are still scanned or photographed into image files for storage.Complex document images contain a myriad of elements,such as text,charts,tables,formulas,seals,columns,headers,footers,irregular character layouts,and intricate backgrounds.To effectively convert these images into editable electronic documents,it is imperative to not only recognize the character content but also comprehend and reconstruct the structural information.In this thesis,we address the challenges associated with complex document image reconstruction by focusing on both dataset and algorithmic aspects.Our primary contributions can be summarized in four key points:1.To mitigate the issues of small dataset sizes and high manual annotation costs in deep learning-based table structure recognition models,we propose an automatic heterogeneous table annotation method called TableRobot.Our experiments demonstrate that this method achieves an annotation accuracy of 93.2%,effectively expanding the table structure annotation dataset and significantly improving the performance of table structure recognition models.2.We design a tailored graph neural network and data annotation processing method for table detection in challenging scenarios,such as those with missing borders,perspective distortions,and twisted tables.Our method outperforms other image detection algorithms on the ICDAR table competition dataset and achieves an average F1 score that is 13%higher than CNN-based methods on the ICDAR+dataset with various interferences.3.We present CarveNet,an algorithm for irregular text recognition in complex backgrounds that addresses the limitations of existing methods affected by background text and distortion.CarveNet exhibits stateof-the-art performance on both regular and irregular scene text recognition benchmark datasets.Compared to prior work,CarveNet achieves higher accuracy on multiple benchmark datasets for regular and irregular scene text recognition.In comparison to previous work,CarveNet improves the accuracy on irregular datasets SVTP and CT80 by 1.8%and 2.6%,respectively.4.To overcome the shortcomings of existing seal removal techniques,we introduce a two-stage SealErase seal removal network.By manually annotating 80 real seal document images and designing a highfidelity seal synthesis method,we construct a seal dataset containing 8,000 samples.Experimental results reveal that the SealErase network surpasses existing methods in seal removal and background restoration.Additionally,SealErase processing enhances the character recognition accuracy of characters covered by seals by 38.86%.Through these advancements,this thesis contributes to the development of sophisticated methods for the reconstruction of complex document images,paving the way for more accurate and efficient electronic document conversion and analysis.
Keywords/Search Tags:Document Reconstruction, Table Recognition, Table Detection, Seal Removal, Text Recognition
Related items