JJobsSonar

Senior Site Reliability Engineer

jobgether · US

Accountabilities: Design, build, and scale Kubernetes-based infrastructure supporting secure, multi-tenant, and highly available applications. Develop and operate AI tooling infrastructure, including secure AI access patterns, MCP servers, and governance frameworks for production environments. Optimize CI/CD pipelines to improve deployment speed, reliability, automation, and rollback safety. Implement progressive delivery practices such as blue/green deployments and canary releases. Advance Infrastructure as Code practices using tools such as Terraform, Helm, and GitOps workflows to create reusable infrastructure patterns. Operate and improve streaming and analytics infrastructure, including Kafka, Flink, and ClickHouse environments. Establish and enhance observability practices through monitoring, SLOs, alerting systems, and operational dashboards. Lead incident response activities, perform root cause analysis, and drive long-term reliability improvements. Build automated testing practices into the software delivery lifecycle. Mentor engineers and promote best practices across Kubernetes, cloud infrastructure, automation, and reliability engineering. Requirements: 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related roles with significant production Kubernetes experience. Hands-on experience integrating AI/LLM tools into engineering or operational workflows, including understanding security, governance, and access control considerations. Proven experience designing and maintaining CI/CD pipelines using tools such as GitHub Actions, Jenkins, GitLab CI, or similar technologies. Strong knowledge of Kubernetes internals and managed cloud Kubernetes services such as EKS, GKE, or AKS. Experience with Infrastructure as Code tools including Terraform, Helm, Pulumi, or equivalent solutions. Proficiency in scripting or programming languages such as Python, Bash, or Go. Experience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry. Production experience working with distributed systems, streaming technologies, and analytics platforms such as Kafka, Flink, and ClickHouse. Strong understanding of cloud infrastructure, automation, system reliability, and operational excellence. Excellent communication and collaboration skills with the ability to work effectively across engineering teams. Experience with multi-region Kubernetes environments, chaos engineering, security automation, policy-as-code, or MLOps workflows is a plus. Benefits: Competitive compensation package ranging from BRL 422,500 – BRL 485,000 total compensation (base salary plus bonus). Stock options and equity opportunities. Health benefits and country-specific employee support programs. Unlimited paid time off and flexible leave policies. Paid parental leave. Tuition reimbursement and learning and development opportunities. Flexible remote working environment. Additional employee benefits designed to support professional growth and well-being.
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago