End-to-end GPU Software Integration On SLURM Cluster (HPC-infrastructure)
Home Case Studies End-to-end GPU Software Integration On SLURM Cluster (HPC-infrastructure)
Project Scope
- Deploy GPU-accelerated VASP on OpenHPC SLURM cluster
- Integrate NVHPC compiler and OpenMPI GPU stack
- Enable NVIDIA L40 GPU offloading (cc89 architecture)
- Establish reproducible build workflow
- Provide centralized shared installation for multi-user access
- Ensure CUDA-driver compatibility and runtime stability
Business Challenges
- MPI wrapper incorrectly linked to GNU (gfortran) instead of NVHPC
- CUDA toolkit mismatch with cluster driver version
- GPU driver not visible on login node (driver version = 0)
- Incorrect GPU architecture flags (cc60/70/80 vs cc89)
- Root vs user environment inconsistencies
- Ensuring SLURM GPU allocation before compilation
ARi’s Solutions
- Loaded correct NVHPC + OpenMPI GPU modules
- Allocated GPU node via SLURM before build
- Updated Makefile with -gpu=cc89 for NVIDIA L40
- Aligned CUDA toolkit with installed driver (CUDA 12.x)
- Installed VASP in shared OpenHPC path with modulefile integration
- Validated execution through SLURM GPU job submission
ARi’s Value Proposition
- Accelerated scientific simulations using GPU offload
- Reduced compute time and improved cluster efficiency
- Reproducible and scalable deployment process
- Multi-user shared environment with controlled access
- Clean integration with OpenHPC module system
- Production-ready GPU workflow for future applications
Download
Case Study
Case Studies
Related
https://www.arigs.com/wp-content/uploads/2026/08/End-To-End-Gpu-Software-Integration-on-Slurm-Cluster-HPC-Infrastructure.pdf