쿠쿠 필터 유사도를 적용한 다중 필터 분산 중복 제거 시스템 설계 및 구현

Design and Implementation of Multiple Filter Distributed Deduplication System Applying Cuckoo Filter Similarity

초록

The need for storage, management, and retrieval techniques for alternative data has emerged as technologies based on data generated from business activities conducted by enterprises have emerged as the key to business success in recent years. Existing big data platform systems must load a large amount of data generated in real time without delay to process unstructured data, which is an alternative data, and efficiently manage storage space by utilizing a deduplication system of different storages when redundant data occurs. In this paper, we propose a multi-layer distributed data deduplication process system using the similarity of the Cuckoo hashing filter technique considering the characteristics of big data. Similarity between virtual machines is applied as Cuckoo hash, individual storage nodes can improve performance with deduplication efficiency, and multi-layer Cuckoo filter is applied to reduce processing time. Experimental results show that the proposed method shortens the processing time by 8.9% and increases the deduplication rate by 10.3%.

키워드

분산중복제거빅 데이터쿠쿠 해시다중계층 쿠쿠 필터소프트웨어 스토리지Distributed DeduplicationBig DataCuckoo HashMultilayer Cuckoo FilterSoftware Storage
제목
쿠쿠 필터 유사도를 적용한 다중 필터 분산 중복 제거 시스템 설계 및 구현
제목 (타언어)
Design and Implementation of Multiple Filter Distributed Deduplication System Applying Cuckoo Filter Similarity
저자
김영아김계희김현주김창근
DOI
10.22156/CS4SMB.2020.10.10.001
발행일
2020-10
저널명
융합정보논문지
10
10
페이지
1 ~ 8