
Staff Site Reliability Engineer
Hamilton Barnes 🌳 · United States
SRE / ReliabilityRemote
About this role
The Staff Site Reliability Engineer will lead incident response, own production health of large GPU fleets, and build observability tools. They will shape how GPU infrastructure is operated at scale and act as the senior reliability voice in customer-facing reviews and partner with product engineering on SLOs and error budgets.
Skills & technologies
Must have
- NVIDIA H100/H200/B200/GB200
- NVLink
- NVSwitch
- InfiniBand
- RoCE
- NVLink
- NCCL
- CUDA
- PyTorch
- Go
- Python
- Rust
- Kubernetes
- Slurm
- HPC
- Linux
- Kernel tuning
- NVIDIA driver
- BPF tooling
Read full description
About the job
Staff Site Reliability Engineer - US Remote
Would you be interested in joining a fast-growing AI infrastructure business, operating as a compute liquidity layer for the global AI labs, data centres, and cloud providers.
They partner with leading organisations to deliver GPU infrastructure at scale, routing workloads across providers, GPU generations, and geographies where it is needed most.
This is a small, senior engineering team where individual judgment has a direct impact on every customer's experience.
The role offers the chance to own reliability end-to-end, and shape how GPU infrastructure is operated at frontier scale. with full autonomy and direct customer impact from day one.
Responsibilities
Lead P0/P1 incident response across the full GPU stack, owning triage, postmortems and systemic fixes.
Own production health of multi-thousand GPU fleets across providers including node lifecycle, firmware rollouts and driver upgrades.
Build and maintain GPU health checks, fabric monitoring, observability and automated remediation tooling.
Define on-call practices including rotations, runbooks, escalation paths and blameless incident reviews.
Act as the senior reliability voice in customer-facing incident reviews, architecture deep-dives and partner with product engineering on SLOs and error budgets.
Required Skills & Experience
Multiple years hands-on building and operating large-scale GPU infrastructure.
Deep expertise with NVIDIA H100/H200/B200/GB200 including NVLink/NVSwitch topology and hardware failure modes.
Production experience with InfiniBand, RoCE and NVLink fabrics alongside NCCL, CUDA and PyTorch distributed training.
Production-grade Go, Python or Rust with strong Kubernetes and/or Slurm/HPC experience.
Expert Linux internals covering kernel tuning, NVIDIA driver/CUDA lifecycle and BPF tooling.
$200K/yr - $250K/yrApply now