JJobsSonar

Site Reliability Engineering (SRE)

Coforge · United States

Remote
About the job Job Title: Site Reliability Engineering (SRE) Key Skills: SRE, DevOps, AWS, Kubernetes, Splunk, Incident Management, Team Leadership, Observability Experience: 5-8 Years’ experience Location: United States We at Coforge is seeking an experienced Site Reliability Engineering (SRE) Team Lead to lead a high-performing Application Support SRE team responsible for ensuring the reliability, availability, scalability, and performance of critical customer-facing applications. The ideal candidate will combine strong technical expertise with people leadership, incident management, operational excellence, and stakeholder management to drive continuous service improvement and reliability initiatives. Key Responsibilities: Lead, mentor, and develop a team of Application Support SREs while supporting career growth and skill development. Manage team capacity planning, performance, staffing, and 24x7 support operations. Act as the senior escalation point for critical incidents and high-severity production issues. Establish and track SRE KPIs, OKRs, SLAs, SLOs, SLIs, and error budgets. Own and improve incident, problem, change, and service readiness management processes. Lead major incident response, stakeholder communication, root cause analysis (RCA), and post-incident reviews. Enhance observability practices using Splunk, OpenTelemetry, AppDynamics, Datadog, and related monitoring tools. Improve dashboards, alerting strategies, telemetry coverage, and operational visibility. Collaborate with Development, Infrastructure, and Architecture teams to build reliability into solutions. Drive automation initiatives, CI/CD improvements, self-healing capabilities, and operational runbooks. Support AWS-hosted applications, Kubernetes environments, MuleSoft APIs, and microservices architectures. Analyze performance issues, logs, and application behavior to support Tier 2/Tier 3 escalations. Required Skills & Experience 5-8+ years of experience in Site Reliability Engineering (SRE), DevOps, Production Support, or Platform Engineering. 2-4+ years of team leadership or people management experience. Strong expertise in AWS cloud environments, microservices, and API-driven architectures. Hands-on experience with Kubernetes and containerized applications. Experience with observability and monitoring platforms such as Splunk, OpenTelemetry, AppDynamics, Datadog, or similar. Strong incident management, problem management, and root cause analysis skills. Knowledge of ITIL frameworks and production support best practices. Excellent communication, stakeholder management, and leadership skills. Preferred Skills: Site Reliability Engineering (SRE) Team Leadership / People Management AWS Cloud Kubernetes Splunk / Observability Tools Incident & Problem Management DevOps & Automation SLO, SLI & Error Budgets Microservices & APIs ITIL Framework Root Cause Analysis (RCA) Stakeholder Management CI/CD Practices Production Support Operations
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago