• DocumentCode
    237378
  • Title

    A Lightweight Virtual Machine Image Deduplication Backup Approach in Cloud Environment

  • Author

    Jiwei Xu ; Wenbo Zhang ; Shiyang Ye ; Jun Wei ; Tao Huang

  • Author_Institution
    Inst. of Software, Univ. of Chinese Acad. of Sci., Beijing, China
  • fYear
    2014
  • fDate
    21-25 July 2014
  • Firstpage
    503
  • Lastpage
    508
  • Abstract
    As most clouds are based on virtualization technology, more and more virtual machine images are created within data centers. Depending on the need of disaster recovery, the storage space used for backup would easily sprawl to a TB or PB level with the growth of images. Unfortunately, different images have a large amount of same data segments. Those duplicated data segments will lead to serious waste of storage resource. Although there is a lot of work focus on deduplication storage and could achieve a good result in removing duplicate copies, they are not very suitable for virtual machine image deduplication in a cloud environment. Because huge resource usage of deduplication operations could lead to serious performance interference to the hosting virtual machines. This paper propose a local deduplication method which can speed up the operation progress of virtual machine image deduplication and reduce the operation time. The method is based on an improved k-means clustering algorithm, which could classify the metadata of backup image to reduce the search space of index lookup and improve the index lookup performance. Experiments show that our approach is robust and effective. It can significantly reduce the performance interference to hosting virtual machine with an acceptable increase in disk space usage.
  • Keywords
    cloud computing; image processing; pattern clustering; virtual machines; backup image; cloud environment; data centers; deduplication operations; deduplication storage; disaster recovery; disk space usage; duplicated data segments; index lookup performance; k-means clustering algorithm; lightweight virtual machine image deduplication backup; resource usage; storage resource; storage space; virtualization technology; Cloud computing; Clustering algorithms; Computer architecture; Image segmentation; Indexes; Sampling methods; Virtual machining; cloud computing; deduplication; virtual machine image; virtualization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computer Software and Applications Conference (COMPSAC), 2014 IEEE 38th Annual
  • Conference_Location
    Vasteras
  • Type

    conf

  • DOI
    10.1109/COMPSAC.2014.73
  • Filename
    6899254