Font Size: a A A

Research On General Chaos Engineering Technology Based On Role Granularity

Posted on:2022-10-23Degree:MasterType:Thesis
Country:ChinaCandidate:X WeiFull Text:PDF
GTID:2480306323978289Subject:Computer software and theory
Abstract/Summary:
With the explosive growth of data from Internet,distributed systems play a vi-tal role in providing Internet users with highly available and highly reliable computing or storage services.However,in the actual operation of such systems,various unpre-dictable emergencies often occur,such as node downtime,network partitions,disk fail-ures,etc.Due to the extremely high complexity of the system itself,unexpected disasters may occur in the face of abnormalities.Chaos engineering is an emerging practical discipline,which creates system ab-normalities through active fault injection,exposes problems in advance,identifies and repairs them,and avoids unavailability of services due to unexpected events in actual operation.The application of chaotic engineering technology in distributed systems can help to verify the stability under faults.However,most of the existing chaotic technolo-gies only focus on fault simulation capabilities and cannot provide fault orchestration.Among the few chaotic technologies that provide orchestration capabilities,one type is only for containerized deployment systems,and the other type only provides services for specific storage systems.None of them can provide failure experiments for data reading,writing or verification,and the granularity of experiments is limited to desig-nated machines or containers.Distributed system developers cannot rely on existing tool technologies for failure experiments,and need to spend extra energy on failure test development.In response to the above problems,this paper proposes a general chaotic technology for distributed systems:1)Abstract a set of description models of different faults and fault occurrence processes in distributed systems,and support operations such as read-ing and writing verification in fault experiments;2)Propose a distributed scheduling model based on role granularity,support monitoring role changes in distributed sys-tems,dynamically migrate faults,and support users to develop plug-in faults or read-ing and writing;3)Corresponding injections are designed for common faults Means to support failure simulation on system machines.This paper develops a general chaos platform based on this technology,and conducts failure experiments on different types of distributed systems.The experimental results show that this technology can provide plug-in customized automated failure experiments based on role granularity,helping distributed system developers to perform Stability verification.
Keywords/Search Tags:Chaos Engineering, Distributed System, Fault Injection, Automated Testing, High Availability
Related items