JJobsSonar

Mid-Level SRE Engineer – Distributed Systems at Scale | Remote

Archer Recruitment · Portugal

SRE / ReliabilityRemote

About this role

The role involves owning reliability across a platform handling millions of real-world user interactions, with a focus on making the system more stable through better standards, smarter automation, deeper observability, and an incident culture that turns every outage into a structural improvement. The SRE will set and enforce reliability standards through SLIs, SLOs, and error budget management, architect and maintain highly available, fault-tolerant infrastructure on GCP, drive Kubernetes operations at scale, build automation and internal tooling, and develop and evolve ML-powered observability, anomaly detection, and alerting pipelines. The role includes full ownership of incident response from first signal to post-mortem close.

Skills & technologies

Must have

  • GCP
  • Kubernetes
  • Python
  • Node.js
  • CI/CD

Mentioned in this posting

Read full description

About the job Mid-Level SRE Engineer – Distributed Systems at Scale | Remote Own reliability across a platform handling millions of real-world user interactions GCP, Kubernetes, ML-driven observability tools that match the ambition Remote working with a clear route into senior technical leadership A globally distributed consumer platform on Google Cloud is scaling fast, and reliability is not keeping up by accident. It is keeping up because the SRE function treats it as an engineering problem, one that demands rigour, automation, and the kind of systems thinking that sees failure modes before they materialise. This is a hire for someone who finds that challenge energising. You will not be handed a stable environment and asked to monitor it. You will be expected to make it more stable through better standards, smarter automation, deeper observability, and an incident culture that turns every outage into a structural improvement rather than a war story. The error budget is your compass. Toil is your enemy. The post-mortem is where learning happens. If that framing resonates, keep reading. There is also something genuinely forward-looking about this team. Machine learning is already being applied to how the platform detects degradation, surfaces anomalies and responds to incidents. This is not experimental; it is in production, and it is evolving. The person joining now will be part of defining how it matures. The role comes with real scope, real ownership, and a technical leadership pathway that is there for engineers who want it, not just promised in an interview. Day to day, you will: Set and enforce reliability standards through SLIs, SLOs and error budget management across critical services Architect and maintain highly available, fault-tolerant infrastructure on GCP Drive Kubernetes operations at scale, including service mesh and advanced orchestration Build automation and internal tooling that permanently removes operational toil Develop and evolve ML-powered observability, anomaly detection and alerting pipelines Take full ownership of incident response from first signal to post-mortem close What the role needs from: SRE or Production Engineering experience in a scaled environment Confident, hands-on GCP or AWS experience running production workloads Real Kubernetes experience is not just familiarity, but operating clusters under pressure Python or Node.js for automation, tooling and scripting A strong understanding of CI/CD and what a good delivery infrastructure looks like What you will get: Competitive base salary plus bonus, healthcare and life assurance Fully remote with the option to work from a shared hub when it suits A technical leadership track with real progression for strong engineers For more information, contact Sam in confidence on +353 1 649 8502 or samer.jaffer@archer.ie
Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago

Senior SRE (Cloud)

Hazelcast · United Kingdom

SRE / ReliabilityRemoteEasy apply1mo ago