JJobsSonar

ML Platform Engineer

Confidential · EMEA

SRE / ReliabilityRemote

About this role

The role involves owning the reliability and scalability of large language model (LLM) serving in production. Responsibilities include building and running deployment, autoscaling, and orchestration on Kubernetes, instrumenting serving with real observability metrics, setting and defending SLOs, and load-testing the platform against traffic spikes. The individual will operate the serving stack (vLLM / Triton / TensorRT-LLM) as a dependable production system. This role is part of a small, senior team at an established enterprise software company building LLM-powered capabilities into its products.

Skills & technologies

Must have

  • Kubernetes
  • cloud infrastructure
  • observability
  • SLOs
  • monitoring
  • incident response
  • GPU-backed workloads
  • software engineering fundamentals

Nice to have

  • serving ML/LLM models
  • inference frameworks
  • performance/load-testing

Read full description

About the job Keep large language models serving reliably at scale — fast, observable, and always up. This is the production-engineering backbone of an AI platform: deployment, autoscaling, observability, and the SLOs that keep inference humming under real traffic. You'll join a small, senior team at an established enterprise software company building LLM-powered capabilities into its products. What you'll do: Own the reliability and scalability of LLM serving in production Build and run deployment, autoscaling, and orchestration on Kubernetes Instrument serving with real observability — time-to-first-token, tokens/sec, latency percentiles, error budgets Set and defend SLOs; load-test and harden the platform against traffic spikes Operate the serving stack (vLLM / Triton / TensorRT-LLM) as a dependable production system What you'll bring: Strong production / SRE / platform-engineering experience running services at scale Deep Kubernetes and cloud infrastructure skills Observability and reliability discipline (SLOs, monitoring, incident response) Comfort operating GPU-backed or ML workloads — a deep ML background is not required Solid software engineering fundamentals Nice to have: Experience serving ML/LLM models specifically Familiarity with inference frameworks (vLLM, Triton, TensorRT-LLM) Performance / load-testing background If you've kept high-scale systems alive and want to move into AI infrastructure, this is a clean on-ramp.
Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago

Senior SRE (Cloud)

Hazelcast · United Kingdom

SRE / ReliabilityRemoteEasy apply1mo ago