
Site Reliability Engineer (43257)
CoolPeople Technology · Czechia
Remote
About the job
I'm looking for a Senior Site Reliability Engineer who will improve observability, monitoring, and automation across a cloud platform while designing scalable telemetry and audit logging solutions. I expect 6+ years of experience in SRE or DevOps, hands-on expertise with Microsoft Azure, Python or Go, ELK Stack, and modern observability practices.
🚀 Project
improving service health monitoring, observability, metrics collection, and alerting across the target platform
collaborating with engineering teams to deliver scalable, operable, and maintainable solutions
implementing automation and engineering improvements to reduce manual operational effort
designing and implementing an end-to-end audit logging and threat detection pipeline using Azure-managed services
developing log preprocessing, filtering, and transformation pipelines prior to data ingestion
managing data ingestion, indexing, and lifecycle policies within Elasticsearch / ELK
integrating telemetry and alerting with enterprise SIEM platforms using Azure Event Hub, Kafka, Cribl, or similar streaming technologies
🎯 Skills
Bachelor's degree or higher in Computer Science, Computer Engineering, or a related field, or equivalent practical experience
6+ years of experience as an SRE, DevOps Engineer, or a similar cloud reliability engineering role supporting production SaaS services
experience working in distributed engineering environments across multiple geographies
hands-on experience with Microsoft Azure
strong programming skills in Python or Go for automation, tooling, or service development
solid knowledge of Unix/Linux internals, systems administration, and networking fundamentals
experience with data streaming platforms such as Kafka, Azure Event Hub, or similar technologies
strong operational knowledge of the ELK Stack (Elasticsearch, Logstash, Kibana)
experience integrating telemetry pipelines with SIEM or enterprise security solutions
experience building observability, monitoring, alerting, metrics, and Operations-as-Code tooling
Ready to apply?Apply now