
(Senior) Site Reliability Engineer - STACKIT Control Plane (m/f/d)
STACKIT · Stadt Heilbronn, Baden-Württemberg, Germany
About this role
The candidate will work at Schwarz Digits to enhance monitoring and alerting infrastructure to shorten time-to-detect and ensure SLOs. Responsibilities include creating playbooks, designing dashboards, and acting as a reliability consultant to development teams. The role involves designing and refining development practices like CI/CD pipelines and progressive delivery strategies. Technical environment involves working with Kubernetes Control Plane internals, Go, Infrastructure as Code, Linux system internals, and distributed datastores and messaging systems. The person will participate in a compensated on-call rotation and lead incident responses and blameless post-mortems.
Skills & technologies
Must have
- Kubernetes
- Go
- Infrastructure as Code
- Linux
- TCP/IP
- CNI
- Load Balancers
- eBPF
- PostgreSQL
- Redis
- Kafka
- NATS
Read full description
Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.
Your tasks
- You collaborate closely with development teams to shorten time-to-detect intervals by enhancing our monitoring and alerting infrastructure and ensuring our services adhere to defined SLOs.
- Your work is critical in continuously optimizing our time-to-mitigation; you achieve this by creating clear playbooks, designing dashboards for first responders, and ensuring our telemetry data (logs and metrics) is comprehensive.
- You act as a reliability consultant to development teams, educating them on reliability patterns and helping them "shift left" to foster a shared responsibility model.
- You design and refine development practices, including CI/CD pipelines, to support progressive delivery strategies such as Canary releases and Blue/Green deployments.
- You proactively analyze and optimize the scalability of the Control Plane, addressing bottlenecks in distributed consensus, database throughput, and kernel-level networking.
- You participate in a compensated on-call rotation, leading incident responses and facilitating blameless post-mortems and Root Cause Analyses.
- You bring 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a specific focus on operating large-scale distributed systems in production.
- You possess expert-level knowledge of Kubernetes Control Plane internals, including the API Server, Controller Manager, Scheduler, and etcd.
- You demonstrate proficiency in Go and write production-grade code to build automation tools, Kubernetes Operators, or glue code that integrates disparate systems.
- You hold deep experience with Infrastructure as Code and container infrastructure, alongside proficiency in Linux system internals (kernel tuning, memory management) and networking (TCP/IP, CNI, Load Balancers, eBPF).
- You bring experience in operating datastores (e.g., PostgreSQL, Redis) and messaging systems (e.g., Kafka, NATS) in scalable environments.
- You run towards fires to learn from them, you automate yourself out of a job, and you believe that hope is not a strategy.