Font Size: a A A

Generative Adversarial Networks Based Image Editing And Multi Modal Image Translation

Posted on:2023-10-21Degree:MasterType:Thesis
Country:ChinaCandidate:J X LiuFull Text:PDF
GTID:2558307154974599Subject:Electronic information
Abstract/Summary:
With the continuous development of Generative Adversarial Nets(GANs)in recent years,GANs have been successfully applied in the field of image editing and multimodal image translation.In the field of image editing,the latent space editing method uses vector operations in the latent space to edit generated images,which can be used to change the specific semantics of the image,but the latent space itself has the problem of semantic entanglement,which leads to latent space editing to change other attributes of the image when editing only one attribute.This restricts the further application of latent space editing.Multi-modal image translation can be used to synthesize missing multi-modal image.The existing methods can receive input from multiple modalities at the same time while the relationship between modalities is ignored,resulting in unsatisfactory image synthesis quality.In response to these problems and challenges,this article has done the following research:(1)Aiming at the semantic entanglement problem in the latent space of GANs,we proposes a generative confrontation network ACGAN based on attribute consistency constraints.The generator input is decomposed into content variables and attribute variables,and the attribute regressor is used in the process of training GANs Constrain the consistency of attribute variables with the attributes of the generated image.The attribute variables are located in the orthogonal attribute space,each element represents the strength of each attribute,and the basis vector of the attribute space is used as a predefined semantic direction.Since the semantic direction is orthogonal,the detanglement between attributes is ensured.Experiments show that ACGAN achieves more thorough attribute semantic detanglement than existing methods,and attribute editing is more accurate.(2)Aiming at the problem of low quality of existing multi-modal image translation,JAGAN is proposed to use the attention mechanism to model the correlation of multiple modalities,fully mine the information of the input modalities,and greatly improve the composite image quality.JAGAN contains two important attention modules:intra-modal attention and inter-modal attention.Intra-modal attention is used to extract personality information related to the target modality,and inter-modal attention is used to complement the information between input modalities.The relationship between the two is that intra-modal attention provides richer information and necessary feature compatibility for inter-modal attention.JAGAN verifies the model’s advancement in performance measure and subjective effects compared with existing methods through quantitative and qualitative experimental analysis in two scenes of human faces and medical images.
Keywords/Search Tags:Generative Adversarial Networks, Image Editing, Multimodal Image Translation, Semantic Entanglement
Related items