A Study Of Feature Extraction Algorithm Of Speech Recognition Confidence Measure

Posted on:2011-03-02

Degree:Master

Type:Thesis

Country:China

Candidate:Y J Guo

Full Text:PDF

GTID:2178360308962235

Subject:Pattern Recognition and Intelligent Systems

Abstract/Summary:

PDF Full Text Request

The large vocabulary continuous speech recognition research has been studied for more than two decades, though significant progress has been made, but there is still a considerable distance from the wide range of applications. In the pursuit of overcoming the deficiencies inside recognition algorithm itself, and improving recognition performance, researchers have gradually introduced the concept of confidence measure, to measure in which degree we could trust the result of speech recognition system. In recent years, speech recognition confidence measure has played a very important role in many applications, including speech error detection and correction, no supervision and semi-supervised training, multi-search technology and corpus selection and verification, etc.Based on different feature-combinations, traditional speech recognition confidence is actually a confidence annotation or classification decisions, with mainly information from decoding messages. However, the current confidence features are still limited to isolated and static, while ignoring the the relationship between the words and their surrounding environment; on the other hand, acoustic features are still dominant, while the experiments show that in speech understanding, human beings depend approximately 30% of the information from the syntax, semantics and other non-acoustic knowledge. Therefore, how to dig out the relationship between words and the environment, to extract the characteristics of syntax and semantics of the word so as to enhance recognition performance of post-processing is a very worthwhile study in the field of feature extraction of confidence measure.For purpose above, in addition to build a traditional baseline system of speech recognition confidence annotation, this paper proposed two new confidence features. The first one is environmental feature, including context, dynamic and the global environment features, which extract more valuable information from the intermediate production of decoding, and provide a more comprehensive description of the relationship between words and the environment from both perspectives of space and time. The second is based on topic similarity of the semantic layer of confidence feature extraction algorithm TSS (Topic Similarity based Semantic confidence feature extraction algorithm), using a new theme Model LDA (Latent Dirichlet Allocation) we could calculated the distribution on theme of first the word in recognition results and then in the context. and distribution similarity between the theme and the word could be figured out as the semantic features of words in context. Experiments show that the two features proposed in this paper deeply excavated valuable decoding information, and, after combined with acoustic features, an significant increase in accuracy of confidence annotation experiment has been seen.

Keywords/Search Tags:

confidence measure, environment feature, latent dirichlet allocation, topic model, semantic

PDF Full Text Request

Related items

1	Research On Text Retrieval Based On Topic Analysis
2	Aurora Image Classification Based On Multi-Feature Latent Dirichlet Allocation
3	News Topic Discovery Research Based On The LDA Model
4	Topic Model Based On Dirichlet Process
5	Research On Topic Modeling Method Based On Semantic Distribution Similarity
6	Analysis Model Of Medical Text And Image Based On LDA And LSA And Its Application
7	Study Of Text Evolution Analysis And Prediction Based On Topic Model
8	Research On Rough Classification Of Academic Papers Based On Topic And Semantic Fingerprint Fusion
9	Research On Text Mining Based On Topic Model
10	Research And Implementation Of Distributed Topic Clustering Technology For Text Flow