JJobsSonar

Site Reliability Engineer

Evlo AI · New York, NY

Remote
About the job About The Role The role is responsible for the reliability, scalability, and performance of large-scale production systems. It sits at the intersection of software engineering and systems engineering, focusing on automating infrastructure management and building self-healing systems. The team works to eliminate manual operations through software development, optimizing containerized deployments, and managing high-throughput traffic across distributed multi-region cloud networks. Key Responsibilities Design, implement, and maintain infrastructure-as-code deployments using Terraform, CloudFormation, or Pulumi across AWS or GCP environments Configure, scale, and secure Kubernetes clusters in production, managing Helm charts, service meshes, and ingress controllers Build and maintain robust observability pipelines using Prometheus, Grafana, Jaeger, and Elasticsearch to establish SLIs, SLOs, and error budgets Develop automated CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins to support continuous deployment and rollback capabilities Participate in an on-call rotation, leading incident response, conducting blameless post-mortems, and writing automation to prevent recurrence of systemic issues Optimize cloud spend, resource utilization, and database performance across PostgreSQL, Redis, and Kafka clusters What We Are Looking For 3–6 years of experience in SRE, DevOps, or Infrastructure Engineering roles managing high-traffic, production cloud architectures Strong proficiency in at least one software development language (Python, Go, or Java) alongside advanced Bash scripting skills Deep technical understanding of Linux internals, TCP/IP networking, DNS, SSL/TLS, and modern web application delivery Hands-on expertise with containerization (Docker) and orchestration (Kubernetes) in cloud production environments Proven experience building infrastructure-as-code and operating distributed SQL/NoSQL databases at scale Bonus: Experience with service mesh technologies (Istio, Linkerd), GitOps paradigms (ArgoCD, Flux), or a BS/MS in Computer Science or a related technical discipline
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago