The following GPU partitions are available on the GPU A100 cluster of system Lise.
Partition name | Nodes | CPU | Main memory (GB) | GPUs per node | GPU hardware | Walltime (hh:mm:ss) | Description |
---|---|---|---|---|---|---|---|
gpu-a100 | 36 | Ice Lake 8360Y | 1000 | 4 | NVIDIA Tesla A100 80GB | 24:00:00 | full node exclusive |
gpu-a100:shared | 5 | 4 | NVIDIA Tesla A100 80GB | shared node access, exclusive use of the requested GPUs | |||
gpu-a100:shared:mig | 1 | 28 (4 x 7) | 1 to 28 1g.10gb A100 MIG slices | shared node access, shared GPU devices via Multi Instance GPU. Each of the four GPUs is logically split into usable seven slices with 10 GB of GPU memory associated to each slice |
See Slurm usage how to pass a 24h walltime limit with job dependencies.
Charge rates
Charge rates for the slurm partitions you find in Accounting.
Examples
$ srun --nodes=2 --gres=gpu:4 --partition=gpu-a100 example_cmd
# Note: The two GPUs may be located on different nodes. $ srun --gpus=2 --partition=gpu-a100:shared example_cmd # Note: Two GPUs on the same node. $ srun --nodes=1 --gres=gpu:2 --partition=gpu-a100:shared example_cmd
$ srun --gpus=1 --partition=gpu-a100:shared:mig example_cmd
Hardware configuration
NHR@ZIB offers access to compute nodes equipped with Nvidia A100 GPUs. The GPU A100 partition consists of two login nodes and 42 compute nodes with the following properties for a single node:
2x Intel Xeon "Ice Lake" Platinum 8360Y (36 cores per socket, 2.4 GHz, 250 W)
- 1 TB RAM (DDR4-3200)
- 4x Nvidia A100 (80GB HBM2, SXM), two attached to each CPU socket
- 7.68 TB NVMe local SSD
- 200 GBit/s InfiniBand Adapter (Mellanox MT28908).
The hardware of the login nodes nodes is similar to those of the A100 GPU compute nodes. Notable exceptions are reduced memory (512 GB instead of 1 TB RAM) and no GPUs (no CUDA drivers) on bgnlogin[1-2].