
Senior Site Reliability Engineer
jobgether · US
Accountabilities: As a Senior Site Reliability Engineer, you will be responsible for designing, maintaining, and improving the infrastructure and operational systems that support a modern software platform. You will help ensure reliability, security, and scalability while partnering with engineering teams to resolve issues and optimize performance. Build, maintain, and support cloud-based environments used throughout the software development lifecycle. Monitor systems and applications, implementing alerting, automation, and self-healing capabilities to improve reliability. Design and implement architectural solutions that support new features, resolve technical challenges, and enhance platform performance. Troubleshoot application and infrastructure issues in collaboration with developers and testing teams. Maintain accurate technical documentation, system diagrams, and operational runbooks. Manage Kubernetes environments, service mesh technologies, and cloud infrastructure components. Support configuration management and deployment processes using tools such as Jenkins, Helm, and Ansible. Improve observability practices through monitoring, logging, and performance analysis. Apply AI and machine learning tools to enhance automation, predictive maintenance, incident management, and operational efficiency. Mentor junior engineers and contribute to improving engineering standards and practices. Requirements: The ideal candidate is a technically skilled and proactive engineer with strong experience in cloud infrastructure, automation, and reliability engineering. You should be comfortable working in complex environments, solving problems independently, and collaborating across teams. Several years of experience in site reliability engineering, DevOps, infrastructure engineering, or a related technical field. Strong experience with Kubernetes, containerized environments, and service mesh management. Hands-on experience with major cloud platforms such as AWS, GCP, Azure, or Oracle Cloud. Solid knowledge of Windows and Linux administration. Experience with relational database administration, including Oracle, SQL Server, or similar technologies. Proficiency with scripting languages such as PowerShell or Bash. Experience with infrastructure automation and configuration management tools including Jenkins, Helm, and Ansible. Familiarity with .NET application management and enterprise software environments. Strong analytical and problem-solving skills with the ability to approach challenges creatively and systematically. Experience using AI/ML tools to improve automation, observability, and operational processes. Excellent communication skills with the ability to collaborate effectively across technical teams. Strong leadership skills, organization, and the ability to manage multiple priorities in a fast-changing environment. A proactive mindset with curiosity and commitment to continuous technical learning. Benefits: Competitive annual salary range of $170,000 - $190,000 USD . Medical, dental, vision, and life/LTD insurance coverage. Company-paid employee health insurance premiums and partial dependent coverage. Flexible work schedules and remote work environment designed to support work-life balance. Opportunities to work on innovative technology and complex engineering challenges. Collaborative culture focused on learning, teamwork, and technical excellence. Regular company-wide communication and engagement initiatives. Opportunities for professional growth and skill development.
Ready to apply?Apply now