Henrik Bengtsson
I live and breathe libre open-access, open-source software. I'm excited about all the wonderful FOSS communities that welcome new scientists and developers on a daily basis. I'm a member of the R Foundation, the Bioconductor Technical Advisory Board, and the director of the R Consortium Infrastructure Steering Committee, and a maintainer of way too many R packages ;)
Formally, Associate Professor at University of California San Francisco (UCSF), United States. Co-director of the UCSF Scientific Software Core and the UCSF Open-Source Program Office (OSPO), member of the Computational Biology Core, Helen Diller Comprehensive Cancer Center (UCSF), and advisory, operational, and user-supporting member of several UCSF high-performance compute (HPC) environments.
GitHub: https://github.com/HenrikBengtsson/, https://github.com/futureverse/
Session
Many users reach a point where analyses, simulations, or data pipelines take too long to run. Parallelization can help, but traditional approaches are often complex, error-prone, and hard to adapt across laptops, servers, the cloud, and high-performance compute (HPC) systems.
This talk introduces new capabilities of the Futureverse ecosystem for parallel and distributed processing in R. It is a popular, decade-old, highly validated framework, which recently has reached a major milestone: the ability to parallelize existing R code with a single declaration, e.g.
y <- map(x, slow_fcn) |> futurize()
cv <- glmnet::cv.glmnet(x, y) |> futurize()
The presentation will show how the futurize() pipe function, which leverages powerful on-the-fly transpiler capabilities of R, can be used to parallelize familiar workflows while preserving the original structure of the code. It will also demonstrate how the same analysis can scale from a notebook to local, remote, or cloud-based compute resources with little or no redesign.
The session is aimed at researchers, data scientists, students, and R programmers who want practical ways to speed up real-world code without becoming parallel-computing specialists, as well as those interested in language design, metaprogramming, and parallel programming. Attendees will learn where this approach works well, which pitfalls to avoid, and how the Futureverse ecosystem helps make parallel R more accessible, scalable, and reproducible within minutes. The ultimate goal is to lower the barriers so that more R users can take advantage of compute resources within reach, including HPC clusters, but which historically required too high a technical skill threshold.