BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//txlf2026//talk//D9AYRJ
BEGIN:VTIMEZONE
TZID:US/Central
BEGIN:STANDARD
DTSTART:20251107T000000
TZNAME:CST
TZOFFSETFROM:-0600
TZOFFSETTO:-0600
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RDATE:20270314T030000
TZNAME:CDT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20261101T020000
TZNAME:CST
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:From Log Files to Live Signals: Rethinking HPC Operations on Linux
  Bare Metal - Borislav Sabotinov
DTSTART;TZID=US/Central:20261107T140000
DTEND;TZID=US/Central:20261107T145000
DTSTAMP:20260930T133615Z
UID:pretalx-txlf2026-D9AYRJ@pretalx.com
DESCRIPTION:A large share of High-Performance Compute (HPC) work runs on L
 inux: the schedulers\, bare metal nodes\, job queues\, and MPI processes. 
 The operating model is usually reactive. An engineer submits a job\, waits
  hours or days\, and reads a log file after the damage is done. By then th
 e compute is spent\, the data is stale\, and the decisions that depended o
 n it are late.\n\nThis talk is a real enterprise case study in flipping th
 at model from reactive to proactive. We built a Linux-native automation la
 yer that connects a chat-based control plane to underlying bare metal HPC 
 infrastructure\, serving several thousand internal engineering users runni
 ng millions of simulations per year. Shell scripts and system-level integr
 ation drive event-driven pipelines and feed data into a control plane that
  gives engineers timely\, actionable signals before compute time is wasted
 .\n\nWe combined pre-submit validation\, to catch the failures that never 
 needed to run\, with runtime inspection of solver state to identify stalle
 d or diverging jobs while they are still cheap to kill. Near real-time ale
 rts moved job status out of log files and into chat where users already wo
 rk. The result is hundreds of thousands of compute hours saved per year\, 
 tens of thousands of engineer hours reclaimed\, and millions of dollars in
  cost savings. \n\nThis talk shows that proactive HPC operations at scale 
 are an automation problem you can already solve on the Linux stack you hav
 e.
LOCATION:Bevo
URL:https://pretalx.com/txlf2026/talk/D9AYRJ/
END:VEVENT
END:VCALENDAR
