JJobsSonar

Senior Site Reliability Engineer - Remote

Akamai Technologies · Responses managed off LinkedIn

SRE / Reliability

About this role

The Senior Site Reliability Engineer will own the end-to-end reliability, scalability, and performance of Akamai's Managed Kubernetes offerings. Responsibilities include developing, maintaining, and automating the control plane and data plane, investigating and resolving production incidents, and providing expert-level support for complex customer-facing issues. The role involves executing and monitoring releases, working with cross-functional teams to benchmark product performance, and implementing proactive measures such as capacity planning and performance tuning. The individual will also build tools to automate analytical workflows and help identify areas for new technology investments. The role requires collaboration with development teams to ensure new features meet reliability and performance goals.

Skills & technologies

Must have

  • Kubernetes
  • Python
  • Golang
  • SQL
  • Ansible
  • Terraform
  • Salt
  • etcd
  • API Server
  • Prometheus
  • Grafana

Read full description

About the job Job Description Join our specialized SRE team, focused on the reliability and scale of Managed Kubernetes! Our mission is to support, maintain, and scale our Managed Kubernetes products. Partner with the best You will work to ensure uninterrupted operations from the beginning to the end of a software’s life cycle for Linode Kubernetes engine ( LKE ) & other Akamai Compute Cloud native offerings. As a Senior Site Reliability Engineer, you will be responsible for: Owning the end-to-end reliability, scalability, and performance of our Managed Kubernetes offerings, including developing, maintaining, and automating the control plane and data plane. This includes investigating and resolving production incidents, implementing preventive measures proactively, and providing expert-level support and troubleshooting for complex customer-facing issues. Besides automating and ensuring system stability, you will be executing and monitoring releases and successfully deploying them. You will work with cross-functional teams to understand and benchmark the performance of our key products and services and influence their evolution. Performing proactive measures such as capacity planning, performance tuning and implementing infrastructure as code. Working closely with development teams to ensure that new features and services are designed and deployed in a way that meets the reliability and performance goals of the organization. Building software tools and systems to automate analytical tasks and workflows to increase efficiency and reliability. Maximize reliability, security, and performance by prioritizing Service Level Objectives (SLOs) and adopting a proactive engineering approach. Helping to identify areas for new technology investments. Do what you love To Be Successful In This Role You Will Have over 5 years of industrial experience and bachelor’s degree in computer science or related engineering field Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) is highly regarded Possess experience with Internet protocols (DNS/HTTP/TLS/TCP) Demonstrate experience with coding in one or more of the following languages (python, golang, SQL) Have experience with configuration management tools such as Ansible, terraform, and Salt. Demonstrate expert, hands-on experience with Managed Kubernetes control plane components (e.g., etcd, API Server, scheduler) and cloud provider integrations. Have experience troubleshooting Unix & Linux issues. Demonstrate experience with observability tools such as Prometheus and Grafana & have some familiarity with web applications. Have excellent communication and organizational skills and be able to articulate technical information in an easy-to-understand manner. Have experience in mentoring junior engineers and leading complex technical initiatives. Learn what makes Akamai a great place to work Connect with us on social and see what life at Akamai is like! We power and protect life online, by solving the toughest challenges, together. At Akamai, we're curious, innovative, collaborative and tenacious. We celebrate diversity of thought and we hold an unwavering belief that we can make a meaningful difference. Our teams use their global perspectives to put customers at the forefront of everything they do, so if you are people-centric, you'll thrive here. Working for you Benefits At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life: Your health Your finances Your family Your time at work Your time pursuing other endeavors Our benefit plan options are designed to meet your individual needs and budget, both today and in the future. About Us Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away. Join us Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you!
Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago