Parallelize Your R Code with a Single Tweak - Easier than Ever Before

Many users reach a point where analyses, simulations, or data pipelines take too long to run. Parallelization can help, but traditional approaches are often complex, error-prone, and hard to adapt across laptops, servers, the cloud, and high-performance compute (HPC) systems.

This talk introduces new capabilities of the Futureverse ecosystem for parallel and distributed processing in R. It is a popular, decade-old, highly validated framework, which recently has reached a major milestone: the ability to parallelize existing R code with a single declaration, e.g.

y <- map(x, slow_fcn) |> futurize()
cv <- glmnet::cv.glmnet(x, y) |> futurize()

The presentation will show how the futurize() pipe function, which leverages powerful on-the-fly transpiler capabilities of R, can be used to parallelize familiar workflows while preserving the original structure of the code. It will also demonstrate how the same analysis can scale from a notebook to local, remote, or cloud-based compute resources with little or no redesign.

The session is aimed at researchers, data scientists, students, and R programmers who want practical ways to speed up real-world code without becoming parallel-computing specialists, as well as those interested in language design, metaprogramming, and parallel programming. Attendees will learn where this approach works well, which pitfalls to avoid, and how the Futureverse ecosystem helps make parallel R more accessible, scalable, and reproducible within minutes. The ultimate goal is to lower the barriers so that more R users can take advantage of compute resources within reach, including HPC clusters, but which historically required too high a technical skill threshold.


Objectives

Many scientific analyses, data pipelines, but also day-to-day workflows in R are computationally intensive, but yet written sequentially (because R is single-process by design). When researchers attempt to speed them up, traditional parallelization methods require them to rewrite significant portions of the codebase, hardcode execution backends, and risk lower code coverage and subtle but critical bugs. Some mistakes and bugs (e.g. random number generation and shared memory access) can even lead to silent, incorrect changes in results, which can have serious consequences in sciences (e.g. clinical trials). Also, after updating a codebase to run parallelly, it often becomes cluttered with code that blurs the original logic and flow, making it harder to maintain. For non-experts, this burden is often too large to handle. Even when attempted, we often do not know what we do not know, which increases the risk for introducing, hard-to-detect mistakes.

The Futureverse (https://www.futureverse.org/) was created to address these problems and remove as much friction as possible for R developers and end-users to scale up their processing without having to become experts in parallel execution or even know the details. Since its inception in 2015, Futureverse has become widely used within the R community where its core future package are among the top-0.6% most downloaded R packages with more than 500 direct reverse dependencies growing 30% annually.

This talk introduces futurize (https://futurize.futureverse.org/), a new package in Futureverse that serves as a universal adapter for parallel and distributed processing in R. By leveraging R's powerful metaprogramming (inspecting and modifying the abstract syntax tree) and on-the-fly transpilation (code-to-code translation) capabilities, the futurize() pipe function allows programmers to parallelize familiar workflows (such as familiar map-reduce calls, domain-specific bootstrap calls, etc.) with a single, minimal addition: piping the R expression into futurize(), e.g. update

y <- purrr::map(x, slow_fcn)

to

y <- purrr::map(x, slow_fcn) |> futurize()

The core Futureverse philosophy is to separate the declaration of what should be parallelized from the how and where on the execution environment. This enables the exact same analysis script to scale seamlessly from a local notebook to an HPC cluster or cloud infrastructure without changing a single line of the core logic. This core design philosophy has been essential for reaching this new milestone.

Audience

I am considering the following to be the target audience of this presentation:

  • Researchers, Data Scientists, and R Programmers who want to speed up real-world code (e.g. simulations, bootstraps, or intensive statistical pipelines) without having to become parallel-computing specialists

  • Educators who wish to introduce parallel programming in R to students, but do not have the room for 2-3 lectures on parallelization strategies

  • Language Experts interested in language design, abstract syntax trees (ASTs), and how transpilation (code-to-code translation) can be used to dynamically re-engineer code execution. This part is of interest not only for R developers, but also those coming from Python and Julia, among other higher-level languages

Prerequisites

  • Attendees only need basic-to-intermediate understanding of R
  • No prior experience with parallel or high-performance computing (HPC) is required
  • Attendees that are not R users, but who are developers in other languages (e.g. Python and Julia), should also attend as they will learn about R's powerful metaprogramming capabilities and how they allow for user-friendly, lightweight transpilation techniques like futurize()

Take homes

By the end of this session, attendees will:

  • Understand how to use Futureverse, notably the futurize() pipe function, to parallelize standard sequential R code instantly
  • Appreciate the paradigm of separating parallel execution configuration from core script logic
  • Gain insight into R's metaprogramming capabilities
  • Learn how to transition an analysis seamlessly from local development to distributed HPC or cloud backends within minutes
  • Hopefully return home and start parallelizing R code that they have planned to scale up but never got around to do it

Talk Outline

  1. Parallelization bottleneck in R (4 mins)
  2. Parallelization with Futureverse (intro, examples, backends) (13 mins)
  3. Under the hood (transpilation) (5 mins)
  4. Future looking (roadmap) (3 mins)
  5. Q&A (5 mins)
Henrik Bengtsson

I live and breathe libre open-access, open-source software. I'm excited about all the wonderful FOSS communities that welcome new scientists and developers on a daily basis. I'm a member of the R Foundation, the Bioconductor Technical Advisory Board, and the director of the R Consortium Infrastructure Steering Committee, and a maintainer of way too many R packages ;)

Formally, Associate Professor at University of California San Francisco (UCSF), United States. Co-director of the UCSF Scientific Software Core and the UCSF Open-Source Program Office (OSPO), member of the Computational Biology Core, Helen Diller Comprehensive Cancer Center (UCSF), and advisory, operational, and user-supporting member of several UCSF high-performance compute (HPC) environments.

GitHub: https://github.com/HenrikBengtsson/, https://github.com/futureverse/