
ML Platform Engineer
Confidential · EMEA
About this role
The role involves owning the reliability and scalability of large language model (LLM) serving in production. Responsibilities include building and running deployment, autoscaling, and orchestration on Kubernetes, instrumenting serving with real observability metrics, setting and defending SLOs, and load-testing the platform against traffic spikes. The individual will operate the serving stack (vLLM / Triton / TensorRT-LLM) as a dependable production system. This role is part of a small, senior team at an established enterprise software company building LLM-powered capabilities into its products.
Skills & technologies
Must have
- Kubernetes
- cloud infrastructure
- observability
- SLOs
- monitoring
- incident response
- GPU-backed workloads
- software engineering fundamentals
Nice to have
- serving ML/LLM models
- inference frameworks
- performance/load-testing