
Sr. Cloud Operations Reliability Engineer (SRE)
Ladders · United States
Remote
About the job
For our client, we are seeking a Sr. Cloud Operations Reliability Engineer (sre) to join the team of a leader in the Healthcare space. This role will lead operational priorities focused on execution quality, efficiency, safety, and measurable performance outcomes. You will collaborate with field leadership, support functions, and cross-functional stakeholders to improve accountability, service delivery, and operational consistency. The position offers the opportunity to influence operational scale, customer outcomes, and organizational performance within a healthcare environment.
Location: Remote - US based candidates only, no visa sponsorship available
Compensation: $120,000 – $150,000 annually
Responsibilities
Own service reliability and operational health, establishing SLOs and monitoring strategies
Lead incident response and post-incident processes, troubleshooting complex production issues
Design and implement automation and operational tooling to enhance production readiness
Build observability solutions, ensuring rapid problem detection and response
Conduct performance analysis and capacity planning for cloud services
Collaborate with development teams to ensure deployment reliability and operational best practices
Support disaster recovery planning and ensure operational readiness documentation is up-to-date
Qualifications
Bachelor's degree in Computer Science, Engineering, Information Systems, or related field
10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or related fields
Extensive hands-on experience with production cloud environments on platforms like GCP or AWS
Proficiency with Infrastructure as Code tools (Terraform, CloudFormation) and version control practices
Strong knowledge of incident management and post-incident review processes
Experience with Kubernetes operations and container orchestration
Familiarity with application performance monitoring and distributed tracing
Benefits
Mentorship opportunities to guide junior engineers
Involvement in shaping operational standards and practices
Opportunity to lead and own critical reliability initiatives
Access to ongoing professional development and technology training
Engagement with cross-functional teams for knowledge sharing
Contribution to business continuity and disaster recovery planning
Our client is an equal opportunity employer. We encourage you to apply even if you don’t meet every qualification—your background could be exactly what this team needs.
Desired Skills and Experience
Google Cloud Platform (GCP), AWS, Terraform, Kubernetes, CI/CD practices, Grafana, Prometheus, Datadog, New Relic, Python, Bash, Go, Infrastructure as Code, Distributed tracing, Application Performance Monitoring (APM), Cloud-native services, Incident management processes
$120K/yr - $150K/yrApply now