JJobsSonar

Senior Site Reliability Engineer

Tenth Revolution Group · London Area, United Kingdom

Remote
About the job Senior Site Reliability Engineer (SRE) – OpenShift | Contract 📍 UK Based / Hybrid – occasional London onsite (rare) 💼 Contract 🔐 SC Clearance desirable (eligibility essential) We’re looking for a Senior Site Reliability Engineer to join a major, highly complex technology programme and help take an established engineering platform to the next stage of its SRE maturity. This isn’t a BAU support role. You’ll join a high-performing engineering team working across a heavily customised OpenShift environment, helping establish the frameworks, tooling and practices that underpin a mature Site Reliability Engineering function. What you’ll be doing... You’ll work across the full SRE landscape, including: Establishing and improving SRE frameworks, practices and processes Defining SLIs, SLOs, KPIs and error budgets Building out monitoring, logging, metrics and observability capabilities Working with Prometheus, Grafana, Thanos and OpenTelemetry Supporting distributed tracing and improving end-to-end platform visibility Improving CI/CD and engineering automation Supporting incident, reliability and service-management processes Infrastructure automation using Terraform Scripting with Python, Bash and/or PowerShell Working across Azure, OpenShift and associated storage/platform services Helping shape how the SRE capability scales and matures longer term You’ll also work directly with technical stakeholders and the end client – presenting progress, challenging ideas and helping make decisions around how the platform evolves. What we’re looking for... We don’t need someone who can simply talk about SRE principles – you need to have applied them. You’ll ideally bring: Strong hands-on Site Reliability Engineering experience A deep understanding of SRE principles and how to implement them in practice Experience establishing or maturing SRE frameworks and ways of working Strong Red Hat OpenShift / Kubernetes experience Prometheus and Grafana, ideally alongside Thanos and/or OpenTelemetry Strong observability, monitoring, logging and metrics experience CI/CD and automation experience Terraform Azure experience Python, Bash and/or PowerShell scripting Experience working within complex enterprise environments Experience within highly regulated, Critical National Infrastructure, government or similarly complex environments would be particularly relevant. The type of person who will succeed... This is a high-performing, technically strong team with a flat structure. You’ll need to arrive with enough experience to contribute quickly while still being comfortable collaborating across different areas. We’re looking for someone who can: Have a strong technical opinion – and back it up Challenge solutions constructively rather than accepting them at face value Work effectively with different personalities and communication styles Build credibility with engineers and senior client stakeholders Share knowledge and help raise engineering standards Adapt quickly rather than staying within a narrow technical silo The platform itself is already highly automated and technically advanced. The opportunity here is to help turn that strong technical foundation into a genuinely mature, scalable SRE capability. If you’re an experienced SRE who enjoys building and improving rather than simply maintaining, I’d be keen to speak.
Up to £62.50/hrApply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago