BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pyconhk2024//speaker//SDEGBS
BEGIN:VTIMEZONE
TZID:HKT
BEGIN:STANDARD
DTSTART:20000101T000000
RRULE:FREQ=YEARLY;BYMONTH=1
TZNAME:HKT
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-pyconhk2024-VRPMHX@pretalx.com
DTSTART;TZID=HKT:20241116T140000
DTEND;TZID=HKT:20241116T143000
DESCRIPTION:PySpark is widely adopted for data analysis in distributed comp
 uting environments. It supports not only the standard DataFrame API but al
 so Python User Defined Functions (UDFs)\, Python Data Sources\, Python UDT
 Fs\, and more. However\, debugging and profiling applications in such dist
 ributed environments are often challenging - you can't simply add a breakp
 oint and inspect variables in your IDE.\n\nIn this presentation\, I will d
 emonstrate effective methods for debugging and profiling PySpark applicati
 ons using existing tools. These include profiling tools that utilize cProf
 ile\, a standard Python profiler\, along with various tricks and best prac
 tices for monitoring and debugging PySpark applications.
DTSTAMP:20260710T112438Z
LOCATION:LT7
SUMMARY:How do I debug my PySpark workloads? - Hyukjin Kwon\, Allison Wang
URL:https://pretalx.com/pyconhk2024/talk/VRPMHX/
END:VEVENT
END:VCALENDAR
