Efficient Parquet deduplication and storage

Deduplicating data means keeping only a single occurrence of every piece of data. Applied to cloud storage, it can significantly reduce physical file sizes, reduce storage costs and accelerate I/O. In this talk, I will cover how Data Engineers can use open-source tools to deduplicate Parquet data with Parquet Content Defined Chunking (CDC) and how it works under the hood (minutes 0-10). I'll also present simple examples as well as real-world examples from the 100PB+ of datasets and buckets on Hugging Face (minutes 10-20).


In general, Parquet files share no identical bytes chunks which prevents deduplication. Even Parquet files with nearly identical data don't deduplicate very well. Parquet Content Defined Chunking (CDC) fixes this and enables efficient Parquet deduplication at the lowest level: pages. With CDC enabled, pages can be deduplicated with a content-addressable storage. Outline of the talk:

  • Intro (2min)

  • Open source (3min): I'll present open source tools implementing Parquet CDC and data engines that use it

  • How Parquet CDC works (5min): I'll present the Parquet layout, explain why pages are unstable on slight changes and why this prevents deduplication, and outline the hashing algorithm behind Parquet CDC that makes pages stable

  • Simple examples (5min): showcase deduplication with simple examples: deduplicate Parquet after operations like append, insert, delete and the % of storage saved. I'll be using open source tools for Parquet CDC and for the content-addressable storage.

  • Real world examples (5min): present real-world AI datasets formatted with Parquet CDC from Hugging Face and its content-adressable storage called Xet. I'll show some common operations on AI datasets and the benefits from Parquet CDC.

Quentin Lhoest

Open source developer in AI and head of datasets at Hugging Face. Quentin has been working on open source projects to democratize AI with a focus on data loading and sharing since 2020. His work includes helping research communities and companies in AI for text, audio, image, video and robotics, and also working on compute and storage solutions for AI datasets at Hugging Face.