BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pydata-amsterdam2026//speaker//RZTZ7U
BEGIN:VTIMEZONE
TZID:Europe/Amsterdam
BEGIN:DAYLIGHT
DTSTART:20250911T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251026T030000
RDATE:20261025T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Trillion-Token Pretraining: Building a Foundational Model for paym
 ent data - Martin Iglesias Goyanes\, Raul Souteloquintla
DTSTART;TZID=Europe/Amsterdam:20260911T115000
DTEND;TZID=Europe/Amsterdam:20260911T122000
DTSTAMP:20260809T115155Z
UID:pretalx-pydata-amsterdam2026-EUY789@pretalx.com
DESCRIPTION:Pretraining foundation models on tabular and sequential data p
 resents challenges that differ fundamentally from NLP or vision. This talk
  covers the key design decisions involved in building a payments foundatio
 n model trained on billions of transactions and trillions tokens: tokenisa
 tion of heterogeneous features\, sequence construction\, masking strategie
 s and pretraining objectives for sequential tabular data\, and the archite
 ctural trade-offs between hierarchical and language-model framings. Attend
 ees will leave with transferable techniques for adapting self-supervised p
 retraining to non-text sequential data.
LOCATION:The Grid
URL:https://pretalx.com/pydata-amsterdam2026/talk/EUY789/
END:VEVENT
END:VCALENDAR
