Font Size: a A A

Cross-Modal Generation And Synchronization Identification For Audio-Visual Data

Posted on:2021-02-18Degree:MasterType:Thesis
Country:ChinaCandidate:H D TanFull Text:PDF
GTID:2428330614960349Subject:Signal and Information Processing
Abstract/Summary:
The perception of the outside world by human is based on the acquisition of a variety of information,of which audio-visual information are the main sources.And human beings perform unique capability of comprehension and analysis when processing this multimodal information.Therefore,building similar capability to machine is a key step in the development of artificial intelligence.With the rapid development of information technology,a large amount of audio-visual data has emerged on the Internet.The tremendous amount of information contained in these data provides the possibility for computers to simulate human perception.Hence,under this background,multimodal machine learning has achieved many breakthrough achievements.Audio-visual data appears synchronously most of the time.Therefore,the lack of either will lead to the imperfection and inanimation of the data.However,it is common for the data of one modality to be missing.The cross-modal audio-visual generation system has the ability to generate the missing modality based on the known one.So,it can solve the problem of data loss mentioned above.Most of the existing models use spectrogram as the medium of auditory data to realize the cross-modal generation,but they do not pay special attention to the inherent properties of spectrogram.Therefore,this thesis constructs a cross-modal audio-visual generation system based on Self-Attention mechanism.It focuses on the special structural characteristics of the spectrogram,and introduces Self-Attention mechanism to learn the features of spectrogram at the level of the overall structure.Then,it performs data feature mapping to achieve the crossmodal audio-visual generation.The post-experimental results show that the system performs better than other similar models because of the effect of Self-Attention mechanism.The cross-modal audio-visual generation system requires a large amount of synchronized audio-visual data during training,but it is difficult to get large amounts of such synchronized data.If it is possible to directly determine whether the unknown audio-visual data is synchronized,the synchronized data can be obtained more conveniently.Therefore,this thesis also constructs an audio-visual synchronization identification system.It uses the self-supervised contrastive coding loss to excavate the common correlation information between the naturally synchronized audio-visual data,and then uses the fusion feature as the final decision information to realize the audio-visual synchronization identification.This self-supervised learning method does not rely on the labeled data that requires a lot of manpower and material resources,and has a wilder scope of application under the premise of excellent performance.Specifically,the research results of this thesis are as follows:1.Considering the special structural characteristic of the spectrogram and the advantages of Self-Attention mechanism in simulating the long range and multi-level dependencies in image area,the Self-Attention mechanism is creatively introduced to learn the characteristic information contained in the spectrogram.Better spectrogram feature extraction provides sufficient information for cross-modal audio-visual generation.2.Construct a cross-modal audio-visual generation system,which uses generative adversarial nets as the basic structure and adds the Self-Attention layer to the network.It combines the Hinge loss function and spectral normalization to train the whole system.These operations can improve the performance of the cross-modal audio-visual generation system.3.Construct an audio-visual synchronization identification system.It builds a loss function through the idea of self-supervised contrastive coding to establish the connection between synchronized audio-visual data.It realizes the audio-visual synchronization identification in a self-supervised manner.
Keywords/Search Tags:Audio-visual data, Self-Attention mechanism, Cross-modal generation, Synchronization identification
Related items