BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//euroscipy-2025//speaker//EQSZFK
BEGIN:VTIMEZONE
TZID:Europe/Warsaw
BEGIN:DAYLIGHT
DTSTART:20240821T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20241027T030000
RDATE:20251026T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20250330T030000
RDATE:20260329T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Efficient processing pipelines for large scale molecular datasets 
 in Python - Franciszek Job
DTSTART;TZID=Europe/Warsaw:20250821T110500
DTEND;TZID=Europe/Warsaw:20250821T113500
DTSTAMP:20260913T120417Z
UID:pretalx-euroscipy-2025-RMGD73@pretalx.com
DESCRIPTION:We introduce an extensible Python framework for automated gene
 ration and preprocessing of large-scale chemical datasets. It is based on 
 parallelized and distributed Dask processing for building molecular pipeli
 nes. RDKit\, written in C++ with Python interface\, is leveraged for molec
 ular processing and computation of structural properties. This allows us t
 o process hundreds of millions of molecules on regular-size server units. 
 We also included a suite of analysis scripts for comparing dataset cardina
 lity\, scaffold diversity\, and chemical space metrics. Created software e
 nables efficient pretraining and benchmarking of molecular foundation mode
 ls\, applicable for varying applications in chemoinformatics.
LOCATION:Room 1.20 (Ground Floor\, Shannon)
URL:https://pretalx.com/euroscipy-2025/talk/RMGD73/
END:VEVENT
END:VCALENDAR
