BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//compute-paris-2026//speaker//7UDM3A
BEGIN:VTIMEZONE
TZID:Europe/Paris
BEGIN:STANDARD
DTSTART:20251125T000000
TZNAME:CET
TZOFFSETFROM:+0100
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20261025T030000
RDATE:20271031T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Efficient Parquet deduplication and storage - Quentin Lhoest
DTSTART;TZID=Europe/Paris:20261125T100500
DTEND;TZID=Europe/Paris:20261125T103500
DTSTAMP:20260930T112545Z
UID:pretalx-compute-paris-2026-7MTRRS@pretalx.com
DESCRIPTION:Deduplicating data means keeping only a single occurrence of e
 very piece of data. Applied to cloud storage\, it can significantly reduce
  physical file sizes\, reduce storage costs and accelerate I/O. In this ta
 lk\, I will cover how Data Engineers can use open-source tools to deduplic
 ate Parquet data with Parquet Content Defined Chunking (CDC) and how it wo
 rks under the hood (minutes 0-10). I'll also present simple examples as we
 ll as real-world examples from the 100PB+ of datasets and buckets on Huggi
 ng Face (minutes 10-20).
LOCATION:Room 106
URL:https://pretalx.com/compute-paris-2026/talk/7MTRRS/
END:VEVENT
END:VCALENDAR
