Beyond JSON: How Parquet's New Variant Type Changes Semi-Structured Data
Have you ever put JSON into a string column and thought there HAS to be a better way?
Well now there is! The Parquet community has ratified the Variant type, a native binary encoding for semi-structured data with this one clever trick called shredding that lets columnar storage do what it does best, even on your messiest nested data. Variant stores field names and offsets separately from values or if they show up consistently enough as a regular column. This makes an otherwise unwieldy JSON blob much more easily optimised. You get the speed of a fixed schema for the parts that are predictable, and the flexibility of schemaless for the parts that aren't.
We know that picking column types can sometimes feel overwhelming, so we collected benchmarks for you and we'll show you why projects like Arrow, Spark, Iceberg, Delta Lake and DuckDB are all rushing to adopt it. Stop parsing those messy JSONs and don't store JSON blobs anymore!
The talk opens with a comparison of common approaches for storing semi-structured data in Parquet - string columns, structs, maps - and their tradeoffs. Then we dive into how the Variant type and shredding work under the hood, with benchmarks showing where Variant delivers on its promises.
Started as a physicist, worked as data scientist and engineer, got interested in data tooling and became an Apache Arrow and Parquet contributor, focusing on the C++ and lately Rust implementations. Would like to see numerical computation become more accessible in general purpose languages and frameworks.