
Senior Site Reliability Engineer
RemoteHunter · United States
Remote
About the job
About Our Client:
The organization operates in the cybersecurity ratings industry, providing continuous security assessments for over 12 million companies across 64 countries. It addresses challenges in enterprise cybersecurity by enabling organizations to monitor and manage risks in their digital footprint effectively. The company's patented rating technology supports a broad range of users, including businesses for self-monitoring, third-party risk management, board reporting, and cyber insurance underwriting, contributing to enhanced security resilience on a global scale.
About the Opportunity:
The Senior Site Reliability Engineer will lead the development and optimization of Kubernetes-based infrastructure and CI/CD systems, including the infrastructure supporting AI tooling. This role is pivotal in ensuring production reliability, accelerating engineering delivery, and embedding best practices in automation, observability, and resilience. The position requires collaboration with engineering teams to maintain high availability and security across multi-tenant environments.
Responsibilities:
Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
Build and operate AI tooling infrastructure, including MCP servers and secure AI access patterns for production.
Optimize and maintain CI/CD pipelines to increase reliability and deployment safety.
Implement progressive delivery strategies such as blue/green and canary deployments.
Advance Infrastructure as Code practices using Terraform, Helm, and Argo CD.
Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and ClickHouse.
Integrate automated testing within the CI/CD lifecycle.
Enhance system observability by defining SLOs, alerts, and dashboards.
Lead incident response and postmortems focusing on root cause and durable fixes.
Mentor engineers on Kubernetes, CI/CD, and cloud infrastructure best practices.
Requirements:
Minimum 6 years in SRE, DevOps, or Infrastructure roles with significant Kubernetes experience.
Experience integrating AI/LLM tooling into operational workflows with attention to security and governance.
Proven ability in building and managing CI/CD pipelines using tools such as GitHub Actions, Jenkins, or GitLab CI.
Strong knowledge of Kubernetes internals and managed services (EKS, GKE, or AKS).
Expertise in Infrastructure as Code tools such as Terraform, Helm, or Pulumi and GitOps.
Proficiency in Python, Bash, or Go programming languages.
Familiarity with observability tools like Prometheus, Grafana, Datadog, or OpenTelemetry.
Production experience with Kafka, Flink, and ClickHouse.
Excellent communication and collaboration skills.
Pay Range and Compensation Package:
The estimated total compensation range for this position is $152,000 - $195,000, including base salary and bonus.
Compensation is based on factors including skills, qualifications, experience, and organizational affordability.
Eligibility for annual performance-based incentive awards and equity may be included.
Benefits & Perks:
Competitive salary and stock options specific to each country.
Health benefits.
Unlimited paid time off.
Parental leave.
Tuition reimbursement.
Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.
Note:
RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.
$152K/yr - $195K/yrApply now