| Web information extraction is a currently lovely research fileld, but the mass , isomer and dynamics of web data is a difficult of web information extraction. We can divide web data into two kinds: structural data and unstructured data. We have maturer methords to deal with structural data. However, because traditional database bottom can not deal with unstructured data, a wey that deal with unstructured data need be presented. Many scientists present web matedata in order to slove the problem .Web metadata can transform unstructured data into structural data. It is difficult to construct a metadata standard for web data. This paper construct a Dublin Core metadata for web text data. This kind of metadata can convert web text data which is unstructured data into structurual data.In this paper, we divide Dublin Core metadata into tracing metadata and contental metadata. We fill in tracing metadata by HTML. The mostly research of this paper is filling in contental metadata..(1)On the base of HTML, we can extract DC.title. In order to extract contental metadata we construct matrix model for web text, by which DC.title And DC.creater can be filled in.(2) On the base of matrix model we combine correlational knowledge of faint math to fill in DCsubject and DC.type.(3) Extracting DC.descriotion is a difficult of this paper. In order to fill in DC.description we divide three steps. Firstly, we deal with lengthy sentences by faint similar matrix and form DC.description candidateal sentences WJH1. Secondly, we deal with lengthy paragraph by faint control and form DC.description candidateal sentences WJH2. Lastly, we deal with WJH2 by plane clustering and C_average clustering.(4) The result of experiments show semantic metadata receive a good performance by our information extraction systerm. |