
Senior Site Reliability Engineer
Ladders · United States
Remote
About the job
For our client, we are seeking a Senior Site Reliability Engineer to join the team of a leader in the Information Technology space. This role will lead work at the intersection of data, AI-enabled capabilities, and scalable technology delivery. You will partner with business, technical, and operational stakeholders to build reporting, surface trends, and translate analysis into action. The position offers the opportunity to strengthen data-driven execution and help teams focus on the metrics that matter most within a technology-driven environment.
Location: Remote - US based candidates only, no visa sponsorship available
Compensation: $152,000 – $195,000 annually
Responsibilities
Design, build, and scale Kubernetes infrastructure for multi-tenant applications
Build and operate AI tooling infrastructure and establish secure access for production
Optimize and maintain CI/CD pipelines for reliability and speed
Implement blue/green and canary deployment strategies
Advance Infrastructure as Code practices with Terraform and Helm
Operate and optimize Kafka, Flink, and ClickHouse for analytics
Lead incident responses and conduct postmortems focusing on root causes
Qualifications
6+ years in SRE, DevOps, or Infrastructure roles, specializing in production Kubernetes
Hands-on experience integrating AI/LLM tooling into workflows with security considerations
Proven success building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.)
Strong knowledge of Kubernetes internals and managed services like EKS, GKE, or AKS
Expertise in Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps practices
Proficient in Python, Bash, or Go programming languages
Familiarity with observability tools like Prometheus and Grafana
Benefits
Competitive salary and stock options
Comprehensive health benefits
Unlimited paid time off (PTO)
Parental leave and tuition reimbursement programs
Our client is an equal opportunity employer. We encourage you to apply even if you don’t meet every qualification—your background could be exactly what this team needs.
Desired Skills and Experience
Kubernetes, MCP servers, AI tooling, CI/CD (GitHub Actions, Jenkins, GitLab CI), Infrastructure as Code (Terraform, Helm, Pulumi), Python, Bash, Go, Kafka, Flink, ClickHouse, Observability tools (Prometheus, Grafana, Datadog, OpenTelemetry), Multi-region Kubernetes, Multi-cluster Kubernetes, Chaos engineering, Resilience testing, Security scanning, Compliance automation, Policy-as-code, LLM observability/tracing tools (Langsmith, Langfuse), MLOps workflows, Contributions to open-source Kubernetes or CI/CD projects, Container orchestration best practices Given the role focuses heavily on Kubernetes, knowledge of broader container orchestration principles is likely essential., Cloud infrastructure managementSkills related to managing cloud resources effectively, especially on platforms like AWS, Google Cloud, and Azure, are implied due to Kubernetes' integration in cloud environments., Network security protocolsUnderstanding of security mechanisms that protect containerized applications and AI access is necessary, particularly in a production environment with sensitive data and AI integrations.
$152K/yr - $195K/yrApply now