
Site Reliability Engineer
Haystack · Basingstoke, England, United Kingdom
Remote
About the job
We are seeking a proactive and collaborative individual to join a leading technology firm renowned for developing resilient cloud platforms and making a measurable impact on service reliability. This organisation excels at solving complex operational challenges through innovative engineering and automation.
The Role
Participate in a 24/7 on-call rota for incident management and resolution.
Lead or support major incident response, coordinating with various engineering and product teams.
Develop and improve operational runbooks, conducting post-incident reviews.
Monitor infrastructure health and optimise alerting strategies using SLIs/SLOs.
Automate repetitive operational tasks to enhance efficiency and reduce MTTR.
What You'll Need
Strong Linux systems administration and production environment support experience.
Hands-on expertise with AWS cloud infrastructure, Docker, and Kubernetes.
Scripting/programming skills in Python, Bash, Go, or similar languages.
Solid understanding of networking fundamentals (DNS, TCP/IP, load balancing).
Experience in a 24/7 operations or NOC environment and adeptness in high-pressure situations.
Excellent communication and stakeholder coordination abilities.
What's On Offer
Opportunity to work with cutting-edge cloud native and automation technologies.
Contribute to a culture of continuous improvement and blameless post-mortems.
Play a key role in building resilient and scalable cloud platforms.
Fully remote work option available.
Apply via Haystack today!
Ready to apply?Apply now