Font Size: a A A

Design Of Storage Scheme For Low Storage Overhead Binary Vector Coding Based On Erasure Codes

Posted on:2019-12-30Degree:MasterType:Thesis
Country:ChinaCandidate:J J WangFull Text:PDF
GTID:2428330566461563Subject:Information and Communication Engineering
Abstract/Summary:
With the rapid development of the Internet and the increasingly widespread application of the Internet in modern society,online data has shown a rapid growth trend,and the big data has embarked on the stage.A huge problem has slowly come to people:how to simply and efficiently store and manage amounts of data.Traditional storage methods that put data and management in the same system are already overwhelming in the face of impending mass data.And there are more and more problems,such as:the security performance of the storage system can't be guaranteed,the reliability can't be maintained and the scalability is low.The proposal of a distributed storage system is very good to remedy the defects in this aspect,and it allows massive data to be stored in the network system in a decentralized manner.The proposed method provides great convenience for the storage of massive data at the moment,and has a strong stability.As a result,distributed storage systems have gradually become mainstream storage systems and the range of applications has become increasingly larger.Distributed storage technology,as the name suggests,disperses the data in the system for storage.The realization of this technology mainly uses idle computers and other terminal equipment in the network.At the same time,the stability and security of the storage system are also guaranteed by increasing the storage redundancy in the system nodes.Currently,there are two existing redundancy strategies:a replication-based redundancy strategy and an erasure-based redundancy strategy.If a replication-based storage method is used in a large storage system,since the replication-based redundancy strategy inherently has a great deal of redundancy,if it is applied to a larger storage system,it will increase redundancy and lead to system bloated,performance worse.Because of the large-scale application of large-scale distributed storage systems for mass data generation,the use of erasure-based code-based storage to reduce redundancy has improved system performance,reduced storage costs for storage systems,and improved the reliability of systems.The network-encoded storage scheme is applied to distributed storage to solve the reliability and recoverability problems in distributed storage.The practice and extensive application of network coding in distributed storage solves the problem of future massive data storage.It is of great significance.This paper mainly studies the storage overhead of nodes in distributed storage systems based on erasure codes.The main contents are as follows:1)The CP-ZD(Combination Parameter Zigzag Decodable)code has relatively low coding complexity and relatively small computational overhead,but its storage overhead is relatively large.In order to solve this problem,this paper proposes a distributed storage scheme of single-node,two-packet,low storage overhead binary vector codes,which satisfies the CP-ZD property at the same time.The 2k = n(2<k<8)original data packets are encoded into n encoded data packets.Each node stores one original data packet and one encoded data packet.In this scheme,any k node among the n nodes can recover the original data file.That is,the code satisfies Maximal distance Separable(MDS).This coding scheme satisfies the nature of the CP-ZD code and has a performance advantage over the CP-ZD's storage overhead.2)The single-node 2-pack binary vector code proposed in 1)reduces the storage overhead in the storage system to a certain extent,but there is still room for improvement in theory,and then it is envisaged whether data can be added to each node.In order to reduce the storage cost of the system,a single-node multi-packet binary vector encoding method was proposed.The first is a single-node 3-packet binary vector code scheme.In this scheme,the 2k=n(k= 5,6,7)original data packets are encoded int0 2n check data packets,and then these data packets are stored in the system nodes.Each node stores 1 original data packet and 2 parity packets.At the same time,the coding scheme has the nature of MDS.That is,any node can be randomly selected from the system to recover the original data file.This is followed by a single-node,4-pack binary vector encoding scheme for only 12 distributed nodes.Compared to a single-node 2-packet binary vector encoding scheme,single-node,multi-packet binary vector codes have a smaller storage overhead.And at the same time meet the CP-ZD nature.
Keywords/Search Tags:Distributed storage system, binary vector code, MDS, Zigzag
Related items