
Site Reliability Engineer
Evlo AI · Minneapolis, MN
Remote
About the job
About The Role
The role is responsible for the reliability, scalability, and performance of critical cloud-native production platforms. This position focuses on engineering solutions to prevent operational toil, automating infrastructure provisioning, and architecting resilient systems that support millions of concurrent transactions.
The engineer will collaborate closely with product development teams to foster a modern devops culture, establish SLOs/SLIs, and ensure that software is built with observability, operability, and reliability as core architectural tenets.
Key Responsibilities
Design, provision, and maintain secure, multi-region cloud infrastructure using Terraform and AWS services
Build and optimize container orchestration platforms using Kubernetes (EKS), including ingress controllers, service meshes, and auto-scaling policies
Establish robust observability pipelines using Prometheus, Grafana, Jaeger, and ELK stack to proactively monitor system health and latency
Lead incident response procedures, perform blameless post-mortems, and engineer permanent mitigations for systemic production failures
Develop and maintain automated CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins to support continuous, zero-downtime deployments
Automate repetitive operational tasks (toil) by writing high-quality infrastructure toolings in Go, Python, or Bash
What We Are Looking For
3–6 years of experience in a Site Reliability Engineering, DevOps, or systems engineering role managing high-traffic production environments
Strong hands-on experience with Infrastructure as Code (IaC), specifically Terraform, and cloud platforms like AWS or GCP
Deep expertise in containerization and orchestration, including Docker and production-grade Kubernetes administration
Proficiency in at least one software development language, preferably Go, Python, or Ruby, for systems automation
Solid understanding of networking concepts (DNS, TCP/IP, HTTP/S, VPCs, Load Balancing) and Linux operating system internals
Bonus: Experience with service meshes (Istio/Linkerd), GitOps workflows (ArgoCD), or multi-cloud architecture certifications (AWS Certified Solutions Architect or CKA)
Ready to apply?Apply now