nerdexam
Dell-EMC

E10-001 · Question #148

Which data deduplication method increases the chance of identifying duplicate data even when there is only a minor difference between two documents?

The correct answer is A. Variable-length segment. Data Deduplication Methods File-level deduplication (also called single-instance storage) detects and removes redundant copies of identical files. It enables storing only one copy of the file; the subsequent copies are replaced with a pointer that points to the original file…

Backup, Archive, and Replication

Question

Which data deduplication method increases the chance of identifying duplicate data even when there is only a minor difference between two documents?

Options

  • AVariable-length segment
  • BSingle-instance
  • CFile level
  • DFixed-block

How the community answered

(51 responses)
  • A
    80% (41)
  • B
    4% (2)
  • C
    12% (6)
  • D
    4% (2)

Explanation

Data Deduplication Methods File-level deduplication (also called single-instance storage) detects and removes redundant copies of identical files. It enables storing only one copy of the file; the subsequent copies are replaced with a pointer that points to the original file. File-level deduplication is simple and fast but does not address the problem of duplicate content inside the files. For example, two 10-MB PowerPoint presentations with a difference in just the title page are not considered as duplicate files, and each file will be stored separately. Subfile deduplication breaks the file into smaller chunks and then uses specialized algorithm to detect redundant data within and across the file. As a result, subfile deduplication eliminates duplicate data across files. There are two forms of subfile deduplication: fixedlength block and variable-length segment. The fixed-length block deduplication divides the files into fixed-length blocks and uses a hash algorithm to find the duplicate data. Although simple in design, fixed- length blocks might miss many opportunities to discover redundant data because the block boundary of similar data might be different. Consider the addition of a person's name to a document's title page. This shifts the whole document, and all the blocks appear to have changed, causing the failure of the deduplication method to detect equivalencies. In variable- length segment deduplication, if there is a change in the segment, the boundary for only that segment is adjusted, leaving the remaining segments unchanged. This method vastly improves the ability to find duplicate data segments compared to fixed-block.

Topics

#data deduplication#variable-length segment#deduplication methods#data reduction

Community Discussion

No community discussion yet for this question.

Full E10-001 Practice