BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pyconjp2026//talk//9CQVSR
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
BEGIN:STANDARD
DTSTART:20250822T000000
TZNAME:JST
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Rediscovering DataFrames: 100× Analytics Without Leaving Pandas -
  Auxten Wang
DTSTART;TZID=Asia/Tokyo:20260822T120000
DTEND;TZID=Asia/Tokyo:20260822T124500
DTSTAMP:20260818T083700Z
UID:pretalx-pyconjp2026-9CQVSR@pretalx.com
DESCRIPTION:You've hit the pandas ceiling: 10 GB DataFrames choke\, you ne
 ed vector search for embeddings\, and reading from S3 takes longer than th
 e query itself. The usual upgrades hurt — Spark rewrite\, new ORM\, sepa
 rate vector DB. There's a shorter path: keep the pandas API\, swap the eng
 ine underneath. Trade up to 100× faster analytics\, native vector search\
 , direct querying of S3 / Postgres / 40+ formats as SQL table functions\, 
 and time-series functions. Built on chDB — the in-process build of Click
 House\, which I created and lead at ClickHouse Inc. We'll cover the migrat
 ion cheat-sheet\, concrete latency numbers\, honest tradeoffs vs DuckDB an
 d native pandas\, and close with a live build: a pandas-heavy data pipelin
 e migrated column-by-column. Code published as an open repo.
LOCATION:Dahlia 1
URL:https://pretalx.com/pyconjp2026/talk/9CQVSR/
END:VEVENT
END:VCALENDAR
