JJobsSonar

Site Reliability Engineer

CareerUS Solutions · United States

Remote
About the job Job Summary We are seeking an experienced Site Reliability Engineer (SRE) to join our growing engineering team. The ideal candidate will have a strong background in cloud infrastructure, automation, system reliability, and DevOps practices. You will be responsible for improving the reliability, scalability, and performance of mission-critical applications while driving automation and operational excellence across the organization. Key Responsibilities Design, build, and maintain highly available, scalable, and secure cloud infrastructure. Monitor production systems, identify performance bottlenecks, and implement proactive solutions. Develop automation tools and scripts to reduce manual operational tasks. Manage CI/CD pipelines and streamline deployment processes. Implement Infrastructure as Code (IaC) using tools such as Terraform or CloudFormation. Configure and maintain monitoring, logging, and alerting platforms. Participate in incident response, root cause analysis (RCA), and post-incident reviews. Collaborate with software engineering teams to improve application reliability and performance. Implement disaster recovery, backup, and high-availability strategies. Ensure security best practices are followed across infrastructure and deployments. Optimize system performance, resource utilization, and cloud costs. Required Qualifications Bachelor's degree in Computer Science, Information Technology, or a related field. 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Systems Engineer. Strong experience with cloud platforms such as AWS, Azure, or Google Cloud Platform (GCP). Proficiency with Linux/Unix administration. Hands-on experience with Docker and Kubernetes. Strong knowledge of Terraform, Ansible, or other Infrastructure as Code tools. Experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps. Experience with monitoring tools such as Prometheus, Grafana, Datadog, Splunk, ELK Stack, or New Relic. Strong scripting skills using Python, Bash, or Go. Experience with version control systems such as Git. Strong understanding of networking concepts, DNS, load balancing, firewalls, and TCP/IP. Excellent troubleshooting and problem-solving skills. Preferred Qualifications Experience with microservices architecture. Knowledge of service mesh technologies such as Istio. Experience managing production Kubernetes clusters. Familiarity with SRE principles including SLIs, SLOs, and error budgets. Experience with security best practices and compliance standards. AWS, Azure, Kubernetes (CKA), or Terraform certifications are a plus.
$130K/yr - $165K/yrApply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago