Font Size: a A A

Non-negative Matrix Factorization For Data Representation And Its Clustering Applications

Posted on:2023-12-23Degree:MasterType:Thesis
Country:ChinaCandidate:J WangFull Text:PDF
GTID:2558307100975639Subject:Control Science and Engineering
Abstract/Summary:
With the development of devices such as the internet,data is increasing at an unpredictable rate,especially the increase of image and video data.For the image and video big data with huge volume,diverse changes and high-dimensional complex structure,not only its storage and transmission are facing huge difficulties,but also how to reduce the dimensionality of these high-dimensional data and make full use of the valuable information in them is the most important problem in the field of information processing.Therefore,Data representation of high-dimensional data is hugely challenging.Non-negative matrix factorization is an effective data representation method,which has attracted extensive attentions in the field of data representation due to its strong interpretability and simple solution of optimization model.The researches on dimension reduction and representation of high dimensional data have achieved great successes with the help of non-negative matrix factorization theory.However,these methods also have some disadvantages.Firstly,Most methods focus only on single factor factorization and obtain one clustering solution.However,real data are usually complex and can be described from multiple attributes or sub-features.And,the various attributes provide complementary information of data.Secondly,most of the non-negative matrix factorization methods based on graph regularized allow the learned data representation to maintain a local geometric structure,improving the discriminative ability of the data representation.However,the graph regularized non-negative matrix factorization methods seriously rely on the quality of the predefined graph.the fixed graph relation matrix limits the flexibility of the model.In addition,the adoption of graph Laplacian regularizer also results in high computational complexity of the model,and it is very difficult to deal with large-scale data.Finally,There are usually noises or outliers in real databases,the constructed graph similarity matrix often fails to truly reveal the inherent adjacency structure of data.Based on non-negative matrix factorization method,this thesis mainly conducts the following research work for the above analyzed problems:Firstly,to solve the problem of single factor factorization,we propose an attributed non-negative matrix multi-factorization model.This method simultaneously learns multiple low-dimensional representations of the original data.With Hilbert-Schmitt Independence Criterion term,we explicitly co-regularize different components to enforce the diversity of the jointly learned representations,which effectively captures complementary multi-attribute information.Furthermore,Graph regularization is also introduced to maintain the local structure of the data,effectively improving the discriminative of the data representation.Secondly,in order to solve the problems that graph similarity are often not optimal in the graph regularized non-negative matrix factorization methods and the calculation cost of graph Laplacian regularizer is high,a globality constrained adaptive graph learning non-negative matrix factorization model is proposed.The model adopts self-representation method to learn adaptive global graph similarity matrix,which can more fully capture the intrinsic adjacency structure information and effectively improve the discriminant ability of data representation.On this basis,a graph factorization technique based on cosine distance measure is proposed to help data representation encode the graph structure information and reduce the computational complexity of model.Thirdly,to solve the problem that the global graph adjacency matrix is not sufficient to describe the data adjacency relations,A structured sparse adaptive graph learning non-negative matrix factorization model is proposed.The model proposes the concept of "relational reconstruction" and obtains a relational reconstruction matrix with local structure and global structure of data samples,which is used to construct adaptive graph.This model not only improves the discriminant ability of data representation,but also has strong robustness.
Keywords/Search Tags:Nonnegative Matrix Factorization, Multi-attribute Representation, Graph Laplacian Constraint, Graph Factorization, Relation Reconstruction
Related items