JJobsSonar

Principal Site Reliability Engineer

jobgether · Ireland

Accountabilities: The Principal Site Reliability Engineer will lead initiatives that improve system resilience, operational efficiency, and engineering excellence across the organization. Define and promote Site Reliability Engineering principles, establishing frameworks for reliability, observability, service level indicators (SLIs), service level objectives (SLOs), and error budgets. Drive the adoption of operational excellence practices and ensure reliability metrics are measurable and continuously improved. Design and implement automation solutions that enhance system scalability, reliability, and deployment efficiency. Conduct production readiness assessments and provide architectural guidance to engineering teams to ensure services are built for scale and resilience. Lead initiatives to improve the lifecycle of distributed systems and microservices, from development and deployment to monitoring and optimization. Identify performance bottlenecks, capacity challenges, and operational risks while implementing sustainable solutions. Partner with engineering and product leadership to embed reliability considerations into product development processes. Lead incident management improvements through blameless postmortems, root cause analysis, and systemic remediation initiatives. Mentor engineers and advocate for best practices in reliability engineering, fostering ownership and operational maturity across teams. Contribute to the future vision and strategic direction of the Site Reliability Engineering function. Requirements: The ideal candidate brings deep expertise in distributed systems, reliability engineering, and organizational leadership, combined with a passion for building scalable and resilient platforms. Proven experience designing, operating, and troubleshooting distributed systems and microservices architectures. Strong expertise in observability, monitoring strategies, incident management, and operational excellence frameworks. Demonstrated ability to drive organizational change and influence engineering practices across multiple teams. Extensive experience implementing reliability frameworks, including SLIs, SLOs, error budgets, and production readiness processes. Strong problem-solving skills with a structured and analytical approach to complex technical challenges. Excellent communication and stakeholder management abilities, with experience collaborating across engineering and leadership teams. Experience working with cloud environments, particularly AWS, is highly desirable. Previous exposure to financial services, regulated industries, or mission-critical platforms is considered an advantage. Interest in blockchain technologies, digital assets, or decentralized finance ecosystems is a plus. Master's degree in Computer Science, Engineering, or a related field is advantageous. Benefits: Fully remote opportunity within a globally distributed and collaborative environment. 35 days of paid time off annually, including public holidays. Additional annual leave entitlement based on years of service. Private health insurance coverage. Opportunity to work on innovative technologies and large-scale, high-impact infrastructure projects. Strong emphasis on learning, career progression, and professional growth. Inclusive and diverse workplace culture with employee-led communities and wellbeing initiatives. Exposure to international teams and cross-functional collaboration across multiple regions.
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago