English
Topic 13. Building a Linux compute cluster with the Slurm job scheduler; tuning for high-performance computing
Goal: become familiar with compute cluster architecture and the components of the Slurm scheduler; learn to prepare Ubuntu Server nodes (networking, users, SSH, NFS, time synchronization), install and configure Slurm with MUNGE authentication, and write batch job scripts for sequential, MPI, and hybrid programs, job arrays, and chains of dependent jobs; master node diagnostics, job accounting with sacct, and operating system tuning for high-performance computing.
Lecture contents
- Cluster architecture and nodes — Compute cluster architecture · Preparing the nodes
- SSH and a shared filesystem — SSH and parallel commands · The NFS shared filesystem
- The Slurm scheduler and commands — The Slurm scheduler · User commands
- Job scripts and accounting — Job scripts · Accounting and scheduling
- Configuration, monitoring, and common mistakes — Operating system tuning for high-performance computing · Monitoring and diagnostics · Scalability of an MPI program · Common mistakes