When hardware decides your biology: What we learned from running PCA on an HPC cluster

Principal Component Analysis (PCA) is a standard preprocessing step in genomics data analysis. For large datasets, we run analyses on High-Performance Computing (HPC) clusters and use approximate PCA. In this talk, we will show that different CPU architectures produce different results. We will show code to reproduce the problem and recommend a practical workaround.


Principal Component Analysis (PCA) is the backbone of high-dimensional data analysis. Default settings in single-cell genomics Seurat and Bioconductor/R packages produce different PCA results on different CPU architectures. This affects downstream clustering and may cause rare but meaningful cell clusters to magically disappear, even when we set a random seed like good computational citizens. Similar issues have been reported for scanpy with containerization proposed as a mitigation. We show this is insufficient: even with identical BLAS/LAPACK libraries, CPU-specific instruction sets produce divergent results below the container level.

Using shell scripts and R code, we reproduce the problem on the High-Performance Computing (HPC) cluster of our institution for two use cases: Seurat clustering and Bioconductor data integration. Approximate PCA uses the R package irlba which causes the problem in both cases. Setting approx = FALSE in the function Seurat::RunPCA in Seurat and BSPARAM=BiocSingular::ExactParam() in the function batchelor::FastMnnParam in Bioconductor solves the problem. We further show that this numerical instability propagates through the entire downstream pipeline, with biological consequences: unstable cluster boundaries alter cell population numbers and composition, directly affecting cell annotation and differential expression results.

It may be possible for a seasoned sysadmin to fix this, but not for a biologist trying to define rare cell populations. This PCA behavior is unexpected and has gone unnoticed when scientists still worked on their laptops with fixed hardware. But it shows up as more scientists move their computationally intensive analyses to heterogeneous HPC infrastructure and affects any field relying on iterative approximate PCA. We provide clarifications, reproducible code, and a practical workaround.

Christof Seiler

I am a Principal Investigator in the Center of Experimental Rheumatology at University Hospital Zurich and University of Zurich. I taught statistics at Maastricht University from 2018 to 2025. I received my academic training at Stanford University from 2013 to 2018, and at Inria and University of Bern from 2008 to 2012.