Slinky (Slurm on Kubernetes)¶
Slinky is SchedMD's way of running Slurm on top of Kubernetes: the Slurm control plane runs as Kubernetes pods, and each worker node runs a slurmd pod. You get the normal Slurm experience (sbatch, srun, squeue, sacct) on a Kubernetes-managed cluster.
Slinky is its own orchestrator option when you deploy an instant cluster, alongside plain Kubernetes. If you selected it, this is what you are running. It differs from a conventional Slurm installation in a few user-visible ways, documented in this section.
Architecture¶
graph TD
User([User])
User -->|"SSH :22 admin · :2222 seamless (opt-in)"| Login
subgraph cluster [Cluster private network]
Login["Login node<br/>bastion · Slurm CLI wrappers"]
subgraph service [Service node]
Op["Slinky operator"]
Ctl["slurm-controller<br/>(slurmctld)"]
Acct["accounting<br/>(slurmdbd + DB)"]
Rest["slurm-restapi"]
Pod["login pod"]
end
subgraph workers [GPU workers]
W1["slurmd pod · 8 GPUs<br/>jobs in shared jail"]
Wn["slurmd pod · 8 GPUs<br/>jobs in shared jail"]
end
Login --- service
Ctl --- W1
Ctl --- Wn
end
- The login node is the SSH entry point. The Slurm CLI there is a set of admin wrappers into the login pod; once an admin enables seamless SSH, ordinary users connect over it and land in the login pod as themselves (see User management).
- The whole Slurm management plane (operator, controller, accounting, REST API and login pod) runs on the service node, so workers do nothing but run jobs.
- Each worker runs a
slurmdpod that owns all 8 GPUs. Every job step executes inside a per-job shared jail. /homeis shared across the login node, the login pod and all workers: build environments once, use them from every job.- Storage inside jobs:
/sharedis a common dataset area on the shared filesystem,/localis the worker's NVMe (persists across jobs on that node), and/tmpis a private per-job scratch on the same NVMe, emptied at job end. - The operator owns
slurm.conf. Configuration changes go through the controller resource, not through files; see Changing the Slurm configuration below.
What is different on Slinky¶
- User management: add users with one command,
slinky-user-add, instead of driving Kanidm by hand. Users log in over seamless SSH once it is enabled (--enable-seamless; off by default), and identities can survive cluster re-provisioning via the identity ledger (replay). - The shared jail: every
srun/sbatchstep runs inside a per-job sandbox. Your/home, CUDA and HPC-X are available inside it; GPUs appear only when you request--gpus. - Containers: containers run via Apptainer from inside a step. Pyxis (
srun --container-image) and an in-job Docker daemon are not available. - Observability: Slinky-specific Grafana dashboards, centralized metrics, and custom prolog/epilog hooks.
- Health checks: recurring node, fabric, storage, training, and inference checks; dashboard results; manual runs; and troubleshooting.
What is the same¶
- The Slurm CLI and job model, see the Slurm documentation and Getting started.
- The Lmod/HPC-X modules work (
/opt/hpcxis bind-mounted into the jail). For Python, create a virtualenv under/homewithpython3 -m venv, or run the/usr/local/bin/pytorch.setup.shhelper, which installsuvon first run. - Monitoring access and validation work as on any instant cluster; the dashboards and checks are Slinky-aware.
Changing the Slurm configuration¶
On Slinky clusters there is no /etc/slurm/slurm.conf to edit. The operator generates slurm.conf and distributes it to the controller and all workers, so a hand-edited file is overwritten on the next reconcile. Extra settings go into the controller's extraConf field, whose lines are appended to the generated slurm.conf.
To add or change a setting (example: enable accounting enforcement):
-
Open the controller resource in an editor:
-
Find
spec.extraConf(a multi-line text block) and append your line at the end, keeping the indentation: -
Save and exit. The operator regenerates
slurm.confand applies it live (equivalent toscontrol reconfigure), with no control-plane downtime and no pod restart. -
Verify from the login node:
Two footnotes:
- A small set of Slurm parameters only take effect on a full
slurmctldrestart (noted in the slurm.conf reference). For those, delete the controller pod and let the operator recreate it:kubectl -n slurm delete pod slurm-controller-0. - If you manage the
slurmHelm release yourself, setcontroller.extraConfin the release values (helm upgrade slurm <slurm-chart> -n slurm --reuse-values -f values.yaml) instead: a livekubectl editis overwritten by the nexthelm upgradeof the release.