Font Size: a A A

Adversarial Training For Universal Multimodal Learning

Posted on:2024-04-11Degree:MasterType:Thesis
Country:ChinaCandidate:T A LuoFull Text:PDF
GTID:2568307103474674Subject:Computer Science and Technology
Abstract/Summary:
Multimodal learning endeavors to harness the innate correlations amongst diverse modal data,thereby facilitating a collaborative understanding of multimodal semantics and fostering an enhanced comprehension of complex environments and tasks.As artificial intelligence technology advances,the significance of multimodal learning research grows ever more crucial.Presently,the predominant approach in this field involves generalized multimodal learning,which tackles multiple multimodal tasks within an extensive framework utilizing a "pretraining-finetuning" methodology.Image and textual data constitute the most prevalent data types,with the generic multimodal learning approach examined herein centering on the "image-text" multimodal scenario,encompassing three quintessential tasks: visual question answering,referring expression comprehension,and image-text retrieval.Contemporary generic multimodal learning techniques hinge on multimodal pretraining to attain their objectives,yet face distinct challenges during both pretraining and finetuning phases.In the pretraining stage,multimodal models must glean semantic associations across modalities while concurrently acquiring as fine-grained a semantic alignment as possible.Currently,most methods can only achieve coarse-grained semantic alignment learning.During the finetuning phase,the precipitous reduction in data volume incurs heightened overfitting risks for multimodal models.Some studies have endeavored to incorporate adversarial training to bolster generalization and robustness,although this invariably increases time overhead.To address these concerns,this thesis presents two novel approaches:1.To tackle the fine-grained semantic alignment learning requisite in the pretraining phase of generic multimodal models,this thesis introduces an adversarial masked learning-based generic multimodal pre-training method,LTM.This approach devises an adversarial masked generation network,training it within an adversarial framework alongside a multimodal pretraining model.The adversarial masked generation network endeavors to mask words related to images in text or regions related to text within images,providing a robust foundation for learning cross-modal fine-grained semantic alignment in multimodal models.The conclusive experimental results substantiate the efficacy of this explicitly enhanced cross-modal fine-grained semantic alignment technique.2.To mitigate overfitting during the fine-tuning phase of generic multimodal models,this thesis proposes a swift generic multimodal fine-tuning training methodology,MAP,premised on adversarial feature perturbation.This approach operates within an adversarial training framework,employing an adversarial perturbation generation module to apply adversarial perturbations at the input data feature level.Distinct from extant adversarial training-based multimodal models,this method utilizes a learnable adversarial perturbation generation module to autonomously generate adversarial perturbations based on the multimodal network output.Throughout training,this technique also adopts an expedited training methodology,substantially diminishing the time cost of adversarial training.Ultimately,experimental results across three multimodal tasks corroborate the efficacy and versatility of this adversarial feature perturbation method.
Keywords/Search Tags:Deep Learning, Generic Multimodal Learning, Multimodal Pretraining, Adversarial Training
Related items