GPU Jobs#
Kellogg-affiliated researchers access dedicated GPU nodes through KLC Reserve by submitting SLURM jobs to the kellogg partition with a --gres (generic resource) GPU request.
For CPU vs. GPU concepts and when a GPU helps, see GPU Concepts and Options. For non-GPU batch jobs, see Submitting SLURM Jobs. For host-RAM-bound CPU work on the dedicated high-memory node, see High-Memory Jobs.
Available GPU Nodes#
Node type |
Nodes |
CPUs |
Host RAM |
GPUs per node |
GPU memory |
Typical use |
|---|---|---|---|---|---|---|
L40S |
1 |
64 |
— |
2 |
48 GB per GPU |
LLM inference, rendering, general GPU workloads |
A100 |
1 |
64 |
2 TB |
1 |
80 GB |
Training and inference for medium-to-large models |
H100 |
2 |
64 |
1 TB |
4 |
80 GB per GPU |
Large-scale LLM training/inference, deep learning |
Requesting a GPU with --gres#
Use --gres=gpu:1 unless you need a specific card type. SLURM assigns any available GPU, which usually reduces queue wait time.
Request a specific card when your workload requires it — for example, a model that needs more than 48 GB of GPU memory (ruling out L40S) or code tuned for H100 Tensor Core behavior:
--gres=gpu:1 # any available GPU (preferred default)
--gres=gpu:l40s:1 # specifically an L40S (48 GB)
--gres=gpu:a100:1 # specifically an A100 (80 GB)
--gres=gpu:h100:1 # specifically an H100 (80 GB)
Request multiple GPUs on one node (meaningful on L40S and H100 nodes, which have more than one GPU):
--gres=gpu:h100:4 # all 4 H100s on a single node
Example Batch Script#
#!/bin/bash
#SBATCH --account=kellogg
#SBATCH --partition=kellogg
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1
#SBATCH --gres=gpu:1 # any available GPU
#SBATCH --time=0:30:00
#SBATCH --mem=40G
#SBATCH --output=/kellogg/proj/<your-netid>/slurm-output/slurm-%j.out
module purge all
module use --append /kellogg/software/Modules/modulefiles
module load mamba/24.3.0
source activate /kellogg/proj/<your-netid>/envs/my-gpu-env
python train.py
Submit with sbatch gpu_job.sh.
Key Parameters#
Parameter |
Description |
|---|---|
|
Your Kellogg SLURM account |
|
Routes the job to Kellogg GPU nodes |
|
CPU cores; increase only if your code parallelizes on CPU |
|
Number and type of GPUs |
|
Host RAM (separate from GPU memory) |
|
Maximum wall time; job is killed if exceeded |
|
Log file path ( |
Interactive GPU Session#
For debugging GPU code interactively:
salloc --account=kellogg \
--partition=kellogg \
--nodes=1 \
--ntasks-per-node=1 \
--gres=gpu:1 \
--mem=40G \
--time=01:00:00
When the shell prompt returns, verify the GPU:
nvidia-smi
Release the allocation with exit.
Verify GPU Access in Python#
Before a long run, confirm your environment sees the GPU:
# pytorch_gpu_test.py
import torch
if torch.cuda.is_available():
print(f"CUDA version: {torch.version.cuda}")
print(f"Number of GPUs: {torch.cuda.device_count()}")
print(f"GPU name: {torch.cuda.get_device_name(0)}")
else:
print("CUDA is not available — running on CPU")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using: {device}")
t1 = torch.randn(1000, 1000, device=device)
t2 = torch.randn(1000, 1000, device=device)
result = t1 + t2
print(f"Tensor shape: {result.shape}")
Sample scripts are in the krs-openllm-cookbook GitHub repo .
A video walkthrough of running this in a Jupyter notebook is available in the Quest OnDemand GPU notebook walkthrough .
Check GPU Utilization#
While a job runs, inspect GPU use:
nvidia-smi
watch -n 2 nvidia-smi # refresh every 2 seconds
Low GPU utilization often means the bottleneck is data loading or CPU preprocessing, not the model itself.
LLM Inference on GPU#
For open-source LLM workflows on KLC, see Open Source LLMs on KLC. GPU nodes accelerate larger models and batch inference that are impractical on CPU-only login nodes.
For worked examples that serve a model in Singularity and query it from a login node, see vLLM Inference on KLC Reserve GPUs (OpenAI-compatible API) and Ollama Inference on KLC Reserve GPUs (submit the shared Ollama launcher with sbatch and query from a login node).
For hosted API workflows (OpenAI, Anthropic, etc.), see LLM API Usage.
Multi-GPU Jobs#
Multi-GPU training requires framework support (PyTorch DistributedDataParallel, Hugging Face Accelerate, etc.) and matching --gres requests:
#SBATCH --gres=gpu:h100:4
Ensure your code initializes all visible devices. A single process that only uses cuda:0 leaves other GPUs idle.
Common Failures#
Symptom |
Likely cause |
What to do |
|---|---|---|
|
Job landed on a node without GPU, or driver/CUDA mismatch |
Confirm |
|
Model or batch size exceeds GPU memory |
Reduce batch size; use a smaller model; request A100/H100 instead of L40S |
Job pending indefinitely |
No GPU of requested type free |
Try |
Low |
CPU-bound pipeline |
Profile data loading; increase dataloader workers |
Quest General GPU Access#
This page covers GPU jobs on the kellogg partition — the Kellogg Reserve path. For Northwestern-wide GPUs at scale, use the Quest General Access gengpu partition instead.
Test and debug your GPU workflow on kellogg first. When the workflow is proven and you need more GPUs than Kellogg Reserve can provide, submit to gengpu with your Quest allocation . That path requires a separate Quest allocation, not just a KLC account. See GPU Concepts and Options for the full comparison and GPUs on Quest for submission details.
Further Reading#
Submitting SLURM Jobs — batch scripts, job arrays, monitoring
High-Memory Jobs — host-RAM-bound CPU work on
qhimem0501When to Use KLC Reserve — case studies for GPU workloads