JJobsSonar

Senior Site Reliability Engineer

RemoteHunter · United States

Remote
About the job About Our Client: The organization operates in the digital infrastructure industry, focusing on delivering distributed cloud and edge computing solutions. It addresses challenges in content delivery, security, and AI application scalability by providing a globally distributed platform that supports high-performance and secure digital experiences. About the Opportunity: The Senior Site Reliability Engineer role is responsible for ensuring the uptime, reliability, and scalability of the organization's AI hardware infrastructure. This position plays a critical part in maintaining and optimizing high-density hardware environments, collaborating with product teams to improve system performance and reliability, and driving operational excellence through automation and monitoring. Responsibilities: Develop and scale programmatic tooling and infrastructure-as-code utilities in Python to reduce operational overhead. Integrate automated workflows across ticketing systems to improve incident response times. Utilize AI tools and LLM-assisted development to enhance scripting and system analysis. Work on private cloud and compute technologies to improve availability and latency. Design and implement telemetry pipelines and custom monitoring dashboards. Participate in 24/7 on-call rotations and lead real-time incident management. Coordinate with third-party vendors and on-site technicians to support uptime activities. Requirements: Minimum 5 years of relevant experience with a Bachelor’s degree in Computer Engineering, Computer Science, or equivalent. Proficiency in Python for building operational tools, API integrations, and automation. Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, and Loki. Knowledge of advanced networking including routing, switching, BGP, and IPv4/IPv6. Experience designing service rollouts with operational readiness and alerting criteria. Skilled in developing technical runbooks, leading incident response, and conducting post-mortems. Ability to manage complex technical problems and coordinate cross-functional teams. Pay Range and Compensation Package: For US-based candidates, base salary ranges from $121,400 to $218,600 per year, with compensation determined by experience, skills, certifications, and location. The package may include annual bonuses, equity awards, and an Employee Stock Purchase Plan. Benefits & Perks: Healthcare coverage 401(k) savings plan Paid parental leave Employee assistance program focusing on mental and financial wellness Equal Opportunity Statement: Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin. Note: RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.
401(k) benefitApply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago