JJobsSonar

Staff Site Reliability Engineer

Hamilton Barnes 🌳 · United States

SRE / ReliabilityRemote

About this role

The Staff Site Reliability Engineer will lead incident response, own production health of large GPU fleets, and build observability tools. They will shape how GPU infrastructure is operated at scale and act as the senior reliability voice in customer-facing reviews and partner with product engineering on SLOs and error budgets.

Skills & technologies

Must have

  • NVIDIA H100/H200/B200/GB200
  • NVLink
  • NVSwitch
  • InfiniBand
  • RoCE
  • NVLink
  • NCCL
  • CUDA
  • PyTorch
  • Go
  • Python
  • Rust
  • Kubernetes
  • Slurm
  • HPC
  • Linux
  • Kernel tuning
  • NVIDIA driver
  • BPF tooling

Read full description

About the job Staff Site Reliability Engineer - US Remote Would you be interested in joining a fast-growing AI infrastructure business, operating as a compute liquidity layer for the global AI labs, data centres, and cloud providers. They partner with leading organisations to deliver GPU infrastructure at scale, routing workloads across providers, GPU generations, and geographies where it is needed most. This is a small, senior engineering team where individual judgment has a direct impact on every customer's experience. The role offers the chance to own reliability end-to-end, and shape how GPU infrastructure is operated at frontier scale. with full autonomy and direct customer impact from day one. Responsibilities Lead P0/P1 incident response across the full GPU stack, owning triage, postmortems and systemic fixes. Own production health of multi-thousand GPU fleets across providers including node lifecycle, firmware rollouts and driver upgrades. Build and maintain GPU health checks, fabric monitoring, observability and automated remediation tooling. Define on-call practices including rotations, runbooks, escalation paths and blameless incident reviews. Act as the senior reliability voice in customer-facing incident reviews, architecture deep-dives and partner with product engineering on SLOs and error budgets. Required Skills & Experience Multiple years hands-on building and operating large-scale GPU infrastructure. Deep expertise with NVIDIA H100/H200/B200/GB200 including NVLink/NVSwitch topology and hardware failure modes. Production experience with InfiniBand, RoCE and NVLink fabrics alongside NCCL, CUDA and PyTorch distributed training. Production-grade Go, Python or Rust with strong Kubernetes and/or Slurm/HPC experience. Expert Linux internals covering kernel tuning, NVIDIA driver/CUDA lifecycle and BPF tooling.
$200K/yr - $250K/yrApply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago

Senior SRE (Cloud)

Hazelcast · United Kingdom

SRE / ReliabilityRemoteEasy apply1mo ago