
Site Reliability Engineer
Evlo AI · New York, NY
Remote
About the job
About The Role
The role is responsible for the reliability, scalability, and performance of large-scale production systems. It sits at the intersection of software engineering and systems engineering, focusing on automating infrastructure management and building self-healing systems.
The team works to eliminate manual operations through software development, optimizing containerized deployments, and managing high-throughput traffic across distributed multi-region cloud networks.
Key Responsibilities
Design, implement, and maintain infrastructure-as-code deployments using Terraform, CloudFormation, or Pulumi across AWS or GCP environments
Configure, scale, and secure Kubernetes clusters in production, managing Helm charts, service meshes, and ingress controllers
Build and maintain robust observability pipelines using Prometheus, Grafana, Jaeger, and Elasticsearch to establish SLIs, SLOs, and error budgets
Develop automated CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins to support continuous deployment and rollback capabilities
Participate in an on-call rotation, leading incident response, conducting blameless post-mortems, and writing automation to prevent recurrence of systemic issues
Optimize cloud spend, resource utilization, and database performance across PostgreSQL, Redis, and Kafka clusters
What We Are Looking For
3–6 years of experience in SRE, DevOps, or Infrastructure Engineering roles managing high-traffic, production cloud architectures
Strong proficiency in at least one software development language (Python, Go, or Java) alongside advanced Bash scripting skills
Deep technical understanding of Linux internals, TCP/IP networking, DNS, SSL/TLS, and modern web application delivery
Hands-on expertise with containerization (Docker) and orchestration (Kubernetes) in cloud production environments
Proven experience building infrastructure-as-code and operating distributed SQL/NoSQL databases at scale
Bonus: Experience with service mesh technologies (Istio, Linkerd), GitOps paradigms (ArgoCD, Flux), or a BS/MS in Computer Science or a related technical discipline
Ready to apply?Apply now