Deduplication

Deduplication is a method in which identical data blocks are stored only once and all further occurrences are merely stored as references to them. This reduces storage requirements, costs and transfer volumes without changing the content of the data.

Technically, data is split into blocks and compared via hash values; detected duplicates are replaced by references. This happens either directly when writing (inline) or afterwards (post-process). How strong the effect is depends on the data patterns: backups with many similar backup states and virtual environments with almost identical system images achieve high deduplication rates, whereas already compressed or encrypted data hardly do.

In practice, deduplication is a central lever for cost efficiency – from all-flash systems, whose effective capacity it multiplies, to backup infrastructures, whose retention periods it makes affordable. The method is thus one of the basic tools of efficient Data Management.