Font Size: a A A

A rich metadata filesystem for scientific data

Posted on:2013-03-31Degree:Ph.DType:Dissertation
University:University of Notre DameCandidate:Bui, HoangFull Text:PDF
GTID:1458390008486621Subject:Computer Science
Abstract/Summary:
As scientific research becomes more data intensive, there is an increasing need for scalable, reliable, and high performance storage systems. Such data repositories must provide both data archival services and rich metadata, and cleanly integrate with large scale computing resources. ROARS is a hybrid approach to distributed storage that provides both large, robust, and scalable storage and efficient rich metadata queries for scientific applications. This dissertation presents the design and implementation of ROARS, focusing primarily on the challenge of maintaining data integrity and achieving data scalability. We evaluate the performance of ROARS on a storage cluster compared to the Hadoop distributed file system. We observe that ROARS has read and write performance that scales with the number of storage nodes. We show the ability of ROARS to function correctly through multiple system failures and reconfigurations. We prove that ROARS is reliable not only for daily data access but also for longtime data preservation. We also demonstrate how to integrate ROARS with existing distributed frameworks to drive large scale distributed scientific experiments. ROARS has been in production use for over three years as the primary data repository for a biometrics research lab at the University of Notre Dame.
Keywords/Search Tags:Data, Scientific, ROARS, Storage
Related items