Font Size: a A A

Research On Video-grounded Multi-turn Dialogues Of Question Answering

Posted on:2024-06-06Degree:MasterType:Thesis
Country:ChinaCandidate:H Y ZhangFull Text:PDF
GTID:2568307160459134Subject:Information and Communication Engineering
Abstract/Summary:
With the rapid development of deep learning technology,the deep fusion and understanding of multimodal information have received extensive attention,and Visual Question Answering is one of the challenging tasks.Image question answering or video question answering tasks require deep feature extraction of the input multimodal information and sufficient multimodal interaction to facilitate the generation of an accurate answer.Currently,the main forms of video question answering tasks are single-turn dialogue and selective question answering.Building on this foundation,this thesis explores the forms of multi-turn dialogues and open-ended question answering:for a video and contextual dialogues,an accurate answer with high consistency is generated for a given question by fusing information across multiple modalities.The main work and contributions of this thesis are as follows:1.For video-grounded multi-turn dialogues question answering task,an end-to-end multi-modal transformer network is proposed.In particular,LayerScale regularized spatio-temporal self-attention blocks are first introduced to flexibly perform end-to-end joint training from both video and image data,which effectively avoids traditional offline feature extraction.By using the pre-trained generative language model BART as the backbone architecture,multimodal interaction and dialogue generation can be achieved.This proposed method achieves excellent performance on multiple test benchmarks of the AVSD dataset.2.Aiming at the overfitting problem in the multi-modal network training process,a quality improvement algorithm based on self-distillation is proposed and further expanded into a branch distillation algorithm applied to practical scenarios.The teacher model in self-distillation is obtained by training on the target dataset,and the soft labels provided by the teacher model and the many-to-one inter-layer adaptive learning are used to improve the generalization performance of the student model.The proposed algorithm achieves further performance improvements on multiple test benchmarks of the AVSD dataset.3.In addition,how to extract critical video information using multimodal models to further improve the quality of dialogue generation is explored.On the basis of answer generation task,the auxiliary task of locating relevant video temporal segments is investigated.A localization algorithm based on 2D temporal network is proposed to model video features of different time periods at different moments,and the selfattention mechanism is used to learn the association of features in each time period.Furthermore,combined with the temporal localization module,a multi-task joint training algorithm for answer generation and temporal localization is proposed,and its superiority is verified on the corresponding test benchmark of the AVSD dataset.
Keywords/Search Tags:Multi-turn video question answering, multimodal learning, end-to-end, knowledge distillation, video temporal localization
Related items