JJobsSonar

Site Reliability Engineer

Evlo AI · Minneapolis, MN

Remote
About the job About The Role The role is responsible for the reliability, scalability, and performance of critical cloud-native production platforms. This position focuses on engineering solutions to prevent operational toil, automating infrastructure provisioning, and architecting resilient systems that support millions of concurrent transactions. The engineer will collaborate closely with product development teams to foster a modern devops culture, establish SLOs/SLIs, and ensure that software is built with observability, operability, and reliability as core architectural tenets. Key Responsibilities Design, provision, and maintain secure, multi-region cloud infrastructure using Terraform and AWS services Build and optimize container orchestration platforms using Kubernetes (EKS), including ingress controllers, service meshes, and auto-scaling policies Establish robust observability pipelines using Prometheus, Grafana, Jaeger, and ELK stack to proactively monitor system health and latency Lead incident response procedures, perform blameless post-mortems, and engineer permanent mitigations for systemic production failures Develop and maintain automated CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins to support continuous, zero-downtime deployments Automate repetitive operational tasks (toil) by writing high-quality infrastructure toolings in Go, Python, or Bash What We Are Looking For 3–6 years of experience in a Site Reliability Engineering, DevOps, or systems engineering role managing high-traffic production environments Strong hands-on experience with Infrastructure as Code (IaC), specifically Terraform, and cloud platforms like AWS or GCP Deep expertise in containerization and orchestration, including Docker and production-grade Kubernetes administration Proficiency in at least one software development language, preferably Go, Python, or Ruby, for systems automation Solid understanding of networking concepts (DNS, TCP/IP, HTTP/S, VPCs, Load Balancing) and Linux operating system internals Bonus: Experience with service meshes (Istio/Linkerd), GitOps workflows (ArgoCD), or multi-cloud architecture certifications (AWS Certified Solutions Architect or CKA)
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago