A Comprehensive Study of the Past, Present, and Future of Data Deduplication

Data deduplication, an efficient approach to data reduction, has gained increasing attention and popularity in large-scale storage systems due to the explosive growth of digital data. It eliminates redundant data at the file or subfile level and identifies duplicate content by its cryptographically...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Proceedings of the IEEE 2016-09, Vol.104 (9), p.1681-1710
Hauptverfasser:	Xia, Wen, Jiang, Hong, Feng, Dan, Douglis, Fred, Shilane, Philip, Hua, Yu, Fu, Min, Zhang, Yucheng, Zhou, Yukun
Format:	Artikel
Sprache:	eng
Schlagworte:	Data compression data deduplication Data reduction Data storage systems delta compression Fingerprint recognition Fingerprints Image coding Lists Redundancy Reproduction Security of data Signatures State of the art storage security Storage systems Taxonomy
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Data deduplication, an efficient approach to data reduction, has gained increasing attention and popularity in large-scale storage systems due to the explosive growth of digital data. It eliminates redundant data at the file or subfile level and identifies duplicate content by its cryptographically secure hash signature (i.e., collision-resistant fingerprint), which is shown to be much more computationally efficient than the traditional compression approaches in large-scale storage systems. In this paper, we first review the background and key features of data deduplication, then summarize and classify the state-of-the-art research in data deduplication according to the key workflow of the data deduplication process. The summary and taxonomy of the state of the art on deduplication help identify and understand the most important design considerations for data deduplication systems. In addition, we discuss the main applications and industry trend of data deduplication, and provide a list of the publicly available sources for deduplication research and studies. Finally, we outline the open problems and future research directions facing deduplication-based storage systems.
ISSN:	0018-9219 1558-2256
DOI:	10.1109/JPROC.2016.2571298