| PDF document is applied widely relying on advantages of performance and transmission,and it has become main form of a variety of literature exists on the Internet and an importantresource should be processed by retrieve technology; Therefore, the research on rapididentification the PDF documents containing mathematical expressions is the premise and thefoundation to realize retrieval of mathematical content.Aiming at the mathematical PDF documents with code, a rapid identification methodbased on hierarchical strategy is designed in this paper. First, the content information can beobtained through analysis and information extraction on PDF documents. The correspondingimages of PDF documents can help us get the boundary information of symbols. Then,according to the information we get above, PDF documents can be detected if they containmathematical expressions in terms of geometric level, symbolic content level and combingthe two levels. In the result, we can identify the PDF documents that contain mathematicalexpressions, which lays a foundation for further expressions indexing and matching. Theexperiments show the method can detect the normal PDF documents quickly and effectively. |