| Adaptive Optics(AO)technology compensates for incident distorted wavefront by changing the phase of the wavefront corrector so as to improve the performance of the optical system.As an effective active compensation technology,AO system has achieved great correction results in various fields.However,the traditional closed loop control method regards AO control system as a linear time-invariant system,which makes the traditional control method unable to deal with the uncertainty caused by various errors,and unable to realize the potential of the system to obtain optimal performance.This dissertation finds a combination of traditional AO control methods and deep reinforcement learning and makes an exploratory research to establish a self-learning intelligent control model.The combination of deep learning and reinforcement learning seamlessly connects the perception environment and system control.This AO intelligent control model is versatile and does not depend on establishing an accurate model.It only needs to take interactive learning with the environment,and continuously adjust the control policy taking advantage of the return signal fed back from the outside world and the collected environmental states,can maintain or approach the best performance according to the system state.Specifically,the traditional linear time-invariant control method based on offline modeling can not deal with the following three situations:(1)During the long-term operation of the AO control platform,due to the influence of time-varying factors such as the vibration of the mechanical platform,the alignment error caused by the deviation of the relative position of the wavefront corrector and the wavefront sensor makes the system parameters change abnormally and cannot adapt to the alignment error.(2)The lack of slope information caused by the Hartman sensor’s short of light or the slope measurement error caused by noise.This type of error is directly coupled to the control model,and the transmission of the slope measurement error leads to a decrease or instability in control performance.(3)Time delay is common in AO systems,and the time delay correction error has a great influence on the performance of the system.Therefore,the control method with a static control policy cannot achieve adaptive predictive control.Focusing on the above three situations,this dissertation carries out theoretical analysis and experimental research to establish linear and nonlinearintelligent control models for AO.The model performs online policy optimization based on the current AO environmental characteristics and always meets performance constraint indices,which provides new ideas for solving the problems caused by errors that traditional control methods are difficult to handle such as control performance degradation and system modeling difficulties.The main research contents of this dissertation are as follows:1.The error transmission process of AO system based on the Hartman sensor is inevitable.The error transmission will affect the correction performance of the system,and the maximum compensation or suppression of error transmission can significantly improve the correction performance of the system.The main error sources of AO are divided into five categories:(1)the spatial sampling error caused by the finite division sampling of the wavefront by the H-S lens array;(2)slope measurement error introduced by the noise factor during the slope measurement;(3)slope measurement of H-S sub-aperture is not ideal or the information is missing under strong scintillation conditions;(4)alignment error caused by spatial mismatch between H-S and deformable mirror;(5)time delay correction error caused by system time delay factor.Transforming the above five types of error analysis into the optimization problem of the composite objective function,the gradient information of the composite objective function is derived as an error compensation method,which provides a theoretical basis for the subsequent online learning model based on the gradient information.2.This dissertation proposed a linear learning model of AO system.The model takes the linear combination of the far-field performance index and the square sum of estimation error as a composite objective function,which can adapt to changes in system parameters and does not depend on establishing an accurate model.In order to maintain the good tracking characteristics of the learning model,gradient momentum terms are introduced to accumulate the gradient information of previous iterations,gradually weaken the impact of historical gradient information on the current model training,improve the impact of current gradient information and avoid online sample storage.Meanwhile,the parallel asynchronous optimization method of the model and the initialization policy of the model parameters are also given.The experimental results show that the model takes into account both the compensation of slope information loss and adaptive noise suppression capabilities,which significantly improves the control accuracy of the system,and achieves adaptability under alignment error withoutre-measurement of the response matrix.The model is simple and efficient with certain engineering significance.However,due to the limited learning ability of the linear model,when there is a many-to-one mapping relationship,the learning process is easy to produce linear offset.3.It is modeled based on the problem of the limited learning ability of the linear learning model and the problem of predictive control of turbulent disturbances.A nonlinear dynamic learning model based on deep reinforcement learning theory is proposed,which utilizes the universal mapping property of the neural network to fit the policy function,and achieve the online rolling optimization policy through the deterministic policy gradient method of reinforcement learning.However,when the actual online policy is optimized,the gradient matrix measurement of the objective function of the model is inaccurate or suddenly increases,which may cause the gradient explosion so that the learning model can not work properly.In order to avoid gradient explosion and ensure that the network model converges steadily,the gradient matrix is projected onto a smaller size for cropping and constraints before the gradient is reversed to the network model.At the same time,in order to prevent the learning rate from decaying too fast,it can adapt different learning rates to each network parameter.Three solutions are adopted: first,the use of the history window.Secondly,the use the mean value for the historical window sequence of parameter gradient momentum terms(excluding the current).Thirdly,the final gradient term is the weighted average of the historical window sequence average and the current gradient momentum term.The experimental results show that the nonlinear dynamic learning model has the characteristics of convenient modeling and available online process description,which can make up for the uncertainty caused by model mismatch,distortion,interference and other factors in time.The control accuracy of the model is improved by online error compensation and noise suppression,and its adaptability improves the stability of the system.Because the model can learn the statistical characteristics of turbulence online without offline modeling,and an adaptive predictive control model is realized,which has obvious engineering and theoretical significance. |