Font Size: a A A

Visual Question Answering Research Based On Visual Reasoning And Dense Description

Posted on:2024-03-12Degree:MasterType:Thesis
Country:ChinaCandidate:T HuFull Text:PDF
GTID:2568307115999399Subject:Electronic Information (Computer Technology) (Professional Degree)
Abstract/Summary:
Visual Question Answering,as a current research hotspot,has a wide range of applications in fields such as smart driving,smart wearing,smart home,and assisted medical care.However,existing visual question answering tasks not only require the model to have the ability to infer the visual objects and their relationships contained in the image,but also require deep semantic understanding ability.Recently,visual question answering models based on visual attention methods have demonstrated impressive visual question answering abilities.However,such visual attention methods overlook the multi-dimensional relationship features of complex scenes in images,and lack the ability to construct visual relationships between complex visual regions.At the same time,the existing models do not effectively use the Semantic information implicit in the image,and there is a lack of high-level Semantic information that can not be used for semantic inference.Therefore,visual question answering models not only need to have the ability to infer the visual objects and their relationships contained in images,but also need to have deep semantic understanding abilities.In response to the above issues,this article proposes a series of improvement measures.This paper proposes a question and answer method based on joint reasoning to address the complexity and inefficiency of constructing visual relationships between visual regions.In order to solve the problem of lacking high-level Semantic information in images and unable to make semantic inference,this paper proposes a question answering method based on dense description.In addition,this paper also introduces a gating mechanism to solve the problem of the effectiveness of screening visual information and Semantic information in complex scenes.The main innovative work of this article is as follows:(1)This article proposes a visual question answering model based on gating mechanism for joint relationship reasoning,which constructs visual relationships for high attention regions to enhance the correlation between visual regions.By associating the positional features of different visual regions,visual relationship clues are generated,and self attention mechanism is used to learn the visual relationship features between visual regions.Finally,gating mechanism is used to dynamically select corresponding features to predict answers.(2)This paper proposes a visual question answering model based on dense description and joint relational reasoning,which extracts semantic content from images and associates visual features to obtain high-level Semantic information.This model enriches the overall semantic clues of the model through dense description algorithms,calculates semantic similarity as semantic weights,and obtains the weight features of dense descriptions.The corresponding features are adaptively selected through gating mechanisms to predict answers.The model based on visual reasoning and dense description is evaluated on VQA V1 dataset and VQA V2 dataset.Quantitative evaluation shows that the model has excellent performance compared with the existing models.At the same time,the qualitative analysis also reveals the effectiveness of the model’s prediction answers.
Keywords/Search Tags:Deep learning, Visual Question Answering, Visual reasoning, Dense description
Related items