Deepa
I'm Deepa, a QA Developer at IBM with 4 years testing enterprise storage on Kubernetes and OpenShift. I work on OpenShift Data Foundation, specializing in Disaster Recovery across multi-cluster environments — deploying clusters, validating failovers, and debugging storage issues. Curiosity drives me to explore and experiment. I'm here to share what I've learned along the way.
Session
Scientific computing teams face a high-stakes decision when designing storage infrastructure: which open-source parallel filesystem fits their workloads? With growing diversity — from large-scale simulations and climate modeling to genomics, AI-assisted research, and multi-disciplinary analysis — the wrong choice can mean performance bottlenecks, heavy operational overhead, or unnecessary cost.
This talk offers a grounded comparison of three widely used open-source parallel filesystems:
- Lustre — the long-established leader in traditional HPC, known for high sequential throughput and strong MPI-IO support, typically deployed for high-performance scratch.
- CephFS — the POSIX file interface of Ceph, the unified software-defined storage platform; CephFS shares the same RADOS cluster as Ceph's object (RGW) and block (RBD) interfaces, making it attractive where file, object, and block are needed from one system, particularly in cloud-native and Kubernetes (Rook) environments.
- BeeGFS — a lightweight parallel filesystem valued for simpler deployment and management in mid-to-large scientific clusters.
The session works through the key decision factors:
- Performance characteristics across scientific I/O patterns (sequential vs. mixed).
- Scalability approaches and metadata handling.
- Operational complexity — deployment, monitoring, and maintenance effort.
- Hardware flexibility and integration with modern tooling (Slurm, Rook, Kubernetes).
- Suitability for hybrid architectures — for example, sites pairing Lustre for high-performance scratch with Ceph for home directory, and long-term storage.
Rather than crowning a single winner, the talk gives attendees a clear decision framework — mapping workload type, scale, operational capacity, and hardware constraints to the file storage that best fits — drawn from production usage patterns in scientific computing.