
Site Reliability Engineer
Evlo AI · Denver, CO
Remote
About the job
About The Role
The role focuses on scaling and maintaining the core cloud infrastructure that powers global production services. This position sits at the intersection of software engineering and systems engineering, ensuring that distributed systems are resilient, performant, and highly observable.
The team is responsible for architecting multi-region Kubernetes clusters, defining Infrastructure as Code standards, and building automated CI/CD deployment pipelines. The mission is to eliminate manual intervention through automation, mitigate production incidents, and maintain strict SLAs for millions of concurrent users.
Key Responsibilities
Design, provision, and manage multi-region cloud infrastructure using Terraform, Helm, and AWS/GCP cloud services
Own the availability, latency, performance, efficiency, and capacity management of Kubernetes cluster environments
Implement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and Datadog to proactively detect and diagnose system degradation
Develop internal tooling and automation in Go or Python to simplify deployment workflows and reduce operational toil
Participate in a blameless post-mortem culture and share in an on-call rotation to quickly resolve production incidents
Collaborate with product engineering teams to optimize application performance, containerize services, and architect fault-tolerant distributed systems
What We Are Looking For
3–7 years of experience in SRE, DevOps, or systems engineering roles managing high-traffic production environments
Strong hands-on experience with container orchestration using Kubernetes (EKS, GKE, or self-managed)
Deep proficiency in writing Infrastructure as Code using Terraform or Pulumi
Solid software engineering foundation with strong coding skills in Python, Go, or Bash for automation
Strong understanding of networking concepts (DNS, TCP/IP, VPC peering, load balancing) and Linux systems administration
Bonus: Experience with service meshes (Istio, Linkerd), GitOps workflows (ArgoCD, Flux), or managing relational and NoSQL databases at scale
Ready to apply?Apply now