
Senior Site Reliability Engineer
Tenth Revolution Group · London Area, United Kingdom
Remote
About the job
Senior Site Reliability Engineer (SRE) – OpenShift | Contract
📍 UK Based / Hybrid – occasional London onsite (rare)
💼 Contract
🔐 SC Clearance desirable (eligibility essential)
We’re looking for a Senior Site Reliability Engineer to join a major, highly complex technology programme and help take an established engineering platform to the next stage of its SRE maturity.
This isn’t a BAU support role.
You’ll join a high-performing engineering team working across a heavily customised OpenShift environment, helping establish the frameworks, tooling and practices that underpin a mature Site Reliability Engineering function.
What you’ll be doing...
You’ll work across the full SRE landscape, including:
Establishing and improving SRE frameworks, practices and processes
Defining SLIs, SLOs, KPIs and error budgets
Building out monitoring, logging, metrics and observability capabilities
Working with Prometheus, Grafana, Thanos and OpenTelemetry
Supporting distributed tracing and improving end-to-end platform visibility
Improving CI/CD and engineering automation
Supporting incident, reliability and service-management processes
Infrastructure automation using Terraform
Scripting with Python, Bash and/or PowerShell
Working across Azure, OpenShift and associated storage/platform services
Helping shape how the SRE capability scales and matures longer term
You’ll also work directly with technical stakeholders and the end client – presenting progress, challenging ideas and helping make decisions around how the platform evolves.
What we’re looking for...
We don’t need someone who can simply talk about SRE principles – you need to have applied them.
You’ll ideally bring:
Strong hands-on Site Reliability Engineering experience
A deep understanding of SRE principles and how to implement them in practice
Experience establishing or maturing SRE frameworks and ways of working
Strong Red Hat OpenShift / Kubernetes experience
Prometheus and Grafana, ideally alongside Thanos and/or OpenTelemetry
Strong observability, monitoring, logging and metrics experience
CI/CD and automation experience
Terraform
Azure experience
Python, Bash and/or PowerShell scripting
Experience working within complex enterprise environments
Experience within highly regulated, Critical National Infrastructure, government or similarly complex environments would be particularly relevant.
The type of person who will succeed...
This is a high-performing, technically strong team with a flat structure. You’ll need to arrive with enough experience to contribute quickly while still being comfortable collaborating across different areas.
We’re looking for someone who can:
Have a strong technical opinion – and back it up
Challenge solutions constructively rather than accepting them at face value
Work effectively with different personalities and communication styles
Build credibility with engineers and senior client stakeholders
Share knowledge and help raise engineering standards
Adapt quickly rather than staying within a narrow technical silo
The platform itself is already highly automated and technically advanced. The opportunity here is to help turn that strong technical foundation into a genuinely mature, scalable SRE capability.
If you’re an experienced SRE who enjoys building and improving rather than simply maintaining, I’d be keen to speak.
Up to £62.50/hrApply now