JJobsSonar

GPU Platform Engineer

Lyceum · Zurich, Switzerland

CloudOn-site

About this role

The role involves enabling the expansion of Lyceum’s bare-metal and managed service GPU offerings, with a focus on both customer-facing technical support and the development of a novel software stack. Responsibilities include handling technical customer requests during onboarding and lifecycle management, as well as building and improving managed services like Slurm or Kubernetes. Key performance indicators include response time, error resolution, customer satisfaction, and the development of software offerings on top of bare-metal infrastructure.

Skills & technologies

Must have

  • Kubernetes
  • Slurm
  • NVIDIA software stack
  • NVIDIA drivers
  • CUDA
  • DCGM
  • GPU
  • InfiniBand
  • NVLink
  • RoCE
  • Linux systems administration
  • IP addressing
  • firewalls
  • VPNs
  • load balancing
  • IaC tools
  • Ansible
  • Terraform

Nice to have

  • confidential computing technologies
  • Intel TDX
  • AMD SEV-SNP
  • NVIDIA DGX/HGX systems
  • multi-node GPU clusters
  • MIG (Multi-Instance GPU) partitioning
  • GPU virtualisation
  • HPC
  • AI/ML workloads
  • customer-facing technical role

Read full description

About the job Your mission You will enable the expansion of Lyceum’s bare-metal and managed service GPU offering, taking our EU-sovereign cloud offering to the next step. This exciting role includes both a technical customer-facing function and the opportunity to develop and contribute to a novel software stack – changing the way GPUs are being used today. Your focus: This role has both a customer and a developer component. On the customer side, you’ll be responsible for technical customer requests, both during the onboarding process and later in the lifecycle. On the developer side, you’ll help to build out and improve our managed services such as Slurm or Kubernetes. Your KPIs: Response time / error resolving Customer satisfaction Develop and improve software offering on top of bare metal offering Your profile We consider candidates from diverse backgrounds, with a deep love for technical challenges and the desire to take on ownership beyond what’s reasonably expected. What we’re looking for: 3+ years of experience working with GPU infrastructure (bare metal, VM, cloud) Experience with Kubernetes and/or Slurm cluster deployment and management Hands-on experience with the NVIDIA software stack (drivers, CUDA, DCGM, GPU) Operator Familiarity with high-performance networking (InfiniBand, NVLink, RoCE) Strong Linux systems administration skills Solid networking fundamentals (IP addressing, firewalls, VPNs, load balancing) Familiarity with IaC tools (Ansible, Terraform, or similar) Strong written and verbal communication in English Nice to have: Experience with confidential computing technologies (Intel TDX, AMD SEV-SNP) Experience with NVIDIA DGX/HGX systems or multi-node GPU clusters Knowledge of MIG (Multi-Instance GPU) partitioning and GPU virtualisation Background in HPC or AI/ML workloads Experience in a customer-facing technical role – you can explain complex infrastructure clearly and handle urgent requests professionally
Ready to apply?Apply now

Similar Cloud jobs

All Cloud jobs

Solutions Architect

Haystack · London, England, United Kingdom

CloudRemoteEasy apply1mo ago

Senior Cloud Engineer

Haystack · Vienna, Vienna, Austria

CloudRemoteEasy apply€4,000/month1mo ago

Senior Cloud Engineer

Haystack · Salzburg, Salzburg, Austria

CloudRemoteEasy apply€4,000/month1mo ago

Cloud Engineer

Haystack · Zurich, Zurich, Switzerland

CloudRemoteEasy apply1mo ago

Cloud Architect

emagine · Leiria, Leiria, Portugal

CloudRemoteEasy apply1mo ago