About the job
Mid-Level SRE Engineer – Distributed Systems at Scale | Remote
Own reliability across a platform handling millions of real-world user interactions
GCP, Kubernetes, ML-driven observability tools that match the ambition
Remote working with a clear route into senior technical leadership
A globally distributed consumer platform on Google Cloud is scaling fast, and reliability is not keeping up by accident. It is keeping up because the SRE function treats it as an engineering problem, one that demands rigour, automation, and the kind of systems thinking that sees failure modes before they materialise.
This is a hire for someone who finds that challenge energising. You will not be handed a stable environment and asked to monitor it. You will be expected to make it more stable through better standards, smarter automation, deeper observability, and an incident culture that turns every outage into a structural improvement rather than a war story.
The error budget is your compass. Toil is your enemy. The post-mortem is where learning happens. If that framing resonates, keep reading.
There is also something genuinely forward-looking about this team. Machine learning is already being applied to how the platform detects degradation, surfaces anomalies and responds to incidents. This is not experimental; it is in production, and it is evolving. The person joining now will be part of defining how it matures.
The role comes with real scope, real ownership, and a technical leadership pathway that is there for engineers who want it, not just promised in an interview.
Day to day, you will:
Set and enforce reliability standards through SLIs, SLOs and error budget management across critical services
Architect and maintain highly available, fault-tolerant infrastructure on GCP
Drive Kubernetes operations at scale, including service mesh and advanced orchestration
Build automation and internal tooling that permanently removes operational toil
Develop and evolve ML-powered observability, anomaly detection and alerting pipelines
Take full ownership of incident response from first signal to post-mortem close
What the role needs from:
SRE or Production Engineering experience in a scaled environment
Confident, hands-on GCP or AWS experience running production workloads
Real Kubernetes experience is not just familiarity, but operating clusters under pressure
Python or Node.js for automation, tooling and scripting
A strong understanding of CI/CD and what a good delivery infrastructure looks like
What you will get:
Competitive base salary plus bonus, healthcare and life assurance
Fully remote with the option to work from a shared hub when it suits
A technical leadership track with real progression for strong engineers
For more information, contact Sam in confidence on +353 1 649 8502 or samer.jaffer@archer.ie