Font Size: a A A

Adaptive Cross-modal Image-tactile Signal Reconstruction

Posted on:2023-02-17Degree:MasterType:Thesis
Country:ChinaCandidate:X Y ShiFull Text:PDF
GTID:2568306836968489Subject:Signal and Information Processing
Abstract/Summary:
With the rapid development of mobile communication networks and various smart sensors,there is a growing demand for immersive interactive experiences.Haptics,as an important perceptual modality,is the key to enhance the overall immersion,fluency and engagement of users.However,traditional audio and video transmission schemes cannot cope with the abruptly changing haptic signals and their ultra-high low latency requirements.Therefore,how to design a universal multimodal transmission scheme that can self-adapt to real-time changes of environment and signal during transmission and realize real-time control of data flow to adapt to various remote operation scenarios is an urgent problem to be solved.In addition,with reference to the research results in the field of multimodal deep learning,incorporating the technology of multimodal artificial intelligence into multimodal communication can achieve the effect that is difficult to be achieved by the traditional technology in the field of communication in some scenarios.Based on this,this paper proposes an adaptive cross-modal communication framework,and designs two adaptive cross-modal imagetactile signal reconstruction methods in the framework,and finally demonstrates the effectiveness of the proposed methods through extensive experiments.It mainly contains the following three elements.(1)The design and implementation of the adaptive cross-modal communication framework is completed.Based on both tactile data volume and communication network properties,the transmission method is self-adapted according to the current communication environment.This includes proposing a cross-modal compression scheme to further compress the data volume using complementary modalities without damaging the data quality;secondly,adaptively adjusting the transmission priority according to the application scenario and communication network properties;and finally,using the signal reconstruction model proposed in Chapters 4 and 5 to perform crossmodal signal recovery at the receiving end.In addition,the effectiveness of the transmission framework is verified in a remote communication scenario built by ourselves.(2)An adaptive cross-modal image-tactile signal reconstruction algorithm based on a selfencoder is proposed,which can guarantee transmission services for novel multimodal applications.In the model,effective knowledge is migrated from unpaired visual and haptic domains to paired modal domains to address the limitations of sparse data pairs on model generalization.Meanwhile,the cross-modal semantic consistency is achieved by mining the complementarity and correlation of different modalities.Among them,we combine domain alignment and feature differentiation learning to compensate for the high heterogeneity of inter-modal data,and the shared semantic features of category differentiation can provide a good execution environment for domain alignment.Finally,the effectiveness of the proposed cross-modal reconstruction model is evaluated on the LMT dataset and our own constructed cross-modal dataset.The quality of the final communication experience is also verified in conjunction with the communication framework proposed in Chapter 3.It is easy to see that the fusion of deep learning models and communication frameworks can largely solve the challenges faced by traditional methods.(3)The CGAN-based adaptive cross-modal image-tactile signal reconstruction algorithm is proposed,which can provide high-quality transmission services for novel multimodal applications.The proposed framework contains two modules: a feature extraction module and a signal generation module.In the feature extraction module,an instance weight-based migration enhancement method is utilized to effectively suppress the occurrence of negative migration and thus extract beneficial features.In the signal generation module,two independent conditional generation networks are utilized to connect the extracted features with random signals into the generator,thus achieving crossmodal intergeneration.From the comparison of the results,the output signals of the model proposed in this chapter are significantly improved in quality compared to Chapter 4,but are not as good as the model proposed in Chapter 4 in terms of data category correspondence,which is also caused by the instability of the generative adversarial model.In addition,the effect played by instance weights in migration is also demonstrated explicitly in the experiments.
Keywords/Search Tags:Cross-modal communication, Cross-modal reconstruction, Transfer learning, Autoencoder, CGAN
Related items