
Site Reliability Engineer (43259)
CoolPeople Technology · Czechia
Remote
About the job
I am looking for an experienced Site Reliability Engineer to join a cloud-native platform team focused on automation, reliability, and secure release processes. You will design and implement tooling for automated patch adoption, improve CI/CD validation workflows, and help ensure stable operation of Kubernetes-based services. The role requires strong hands-on experience with Kubernetes, Azure, automation, and production-grade cloud environments.
🚀 Project
automation and reliability improvements for a shared Kubernetes platform
design and implementation of tooling for automated patch adoption
improving release safety through validation and rollout checks in CI/CD pipelines
building validation mechanisms across Kubernetes, Azure, logs, and metrics
integration of patch validation into internal release workflows
support for patching of platform components (external-dns, external-secrets, cert-manager, ingress controllers)
collaboration with engineering teams to ensure scalable, repeatable, and low-friction upgrade processes
contribution to observability, monitoring, and operational simplification
🎯 Skills
6+ years of experience in SRE, DevOps, platform engineering, or cloud infrastructure
strong hands-on experience with Kubernetes and cloud-native tooling
experience with CI/CD, release engineering, and production rollout practices
experience with public cloud (preferably Microsoft Azure)
programming/scripting skills in Python or Go
experience with observability (logs, metrics, monitoring, health checks)
background in maintaining shared platform components and managing lifecycle/patching
💡 Nice to have
experience with platform component upgrades and dependency management
knowledge of patch governance and release validation strategies
experience working in large-scale enterprise cloud environments
familiarity with tools like external-dns, cert-manager, or ingress controllers
Ready to apply?Apply now