Research On The Unsupervised Representation Learning Of Multi-view Data | | Posted on:2023-12-09 | Degree:Doctor | Type:Dissertation | | Country:China | Candidate:L Q Yang | Full Text:PDF | | GTID:1528307169977539 | Subject:Computer Science and Technology | | Abstract/Summary: | | | With the advent of the age of intelligence,a large amount of complex data continues to emerge from various applications,which become the basis for building various machine learning models.Representation learning is one of the core research problems in the field of artificial intelligence,and the quality of complex data representation is directly related to the learning performance of subsequent models.In this paper,we focus on the limitation of unsupervised representation learning,such as the large-scale data,various types of data,diverse of data quality and the distribution shifts.We study representation learning for Euclidean data and non-Euclidean data.We build a series of unsupervised representation learning method and apply them methods to the tasks of clustering analysis and graph classification.The major contributions are as follows:(1)An online binary incomplete multi-view clustering.Firstly,we established the connection between the quality of representation learning and the generalization performance of multi-view clustering.Then,we provide the theoretical analysis of the expected clustering risk of multi-view clustering.Further,we introduce the hash learning into a multi-view clustering framework.To address the challenges of missing views and streaming data in multi-view data,we extend the hash learning-based multi-view clustering framework and propose a binary online clustering algorithm for incomplete multiview data.The design of online optimization algorithms for data streams greatly reduces the space complexity of the algorithm(which is only scale linearly to the block size).Extensive experiments show that the proposed method performed stably under different view incomplete rates and outperforms the state-of-the-art methods.The proposed method is2-13 times faster than the continuous space clustering algorithm in terms of running time.(2)A scalable auto-weighted discrete multi-view clustering method with local structure preservation.We extend the multi-view clustering framework and propose a binary multi-view clustering method(SDMVC)with adaptive weights and local structure preserving.Furthermore,we proposed to learn the parameters using a combination of datadriven and heuristic approaches.The algorithm has linear time and space complexity.Extensive experiments show that the algorithm outperforms the leading benchmark methods for clustering on multiple datasets with a performance improvement of up to 12%.The experiments validate that improving the quality of representation learning helps improve the performance of clustering tasks.(3)A graph self-supervised representation learning method based on adversarial data augmentation.In this paper,a contrastive self-supervised learning method with asymmetric network structure is proposed.In the data augmentation method,the adversarial data augmentation method is designed to reduce the reliance on expert knowledge and manual trial and error.The method is simple and effective.The advantages of the proposed method is that of no maintenance of negative samples and no manual design of data augmentation.Extensive experiments have shown that using adversarial data augmentation on graph classification tasks has a consistent performance improvement of up to 4.7% over hand-designed data augmentation.The proposed method outperforms the state-of-the-art self-supervised learning methods on graph classification tasks.(4)A graph self-supervised representation learning method for out-of-distribution generalization.Since spurious correlations between representations and labels are the main factor causing model performance degradation,this paper proposes to eliminate correlations between representation dimensions in self-supervised representation learning.We convert the elimination of correlation between representation dimensions into seeking the independence of representation dimensions by introducing Hilbert Schmidt independence test.Random Fourier features are used to improve the efficiency of the independence test to accommodate self-supervised training of large-scale data.Extensive experiments in the scenario of distribution shift,the proposed method outperforms other self-supervised learning methods by 5.56%,and the performance improvement is 7.21%on average compared with typical GNN methods. | | Keywords/Search Tags: | representation learning, multi-view clustering, learning to hash, self-supervised learning, graph neural networks, adversarial training | | Related items |
| |
|