JJobsSonar

Site Reliability Engineer

Data Nexus AI · United States

Remote
About the job This is a W2 Role, Please do not apply for C2C We are looking for a highly experienced Senior Observability & Site Reliability Engineer to support large-scale enterprise platforms and mission-critical applications. The ideal candidate will have deep hands-on experience in building and operating end-to-end monitoring, logging, and alerting solutions across distributed environments .This role involves close collaboration with development, infrastructure, and operations teams to ensure platform reliability, performance visibility, and incident response effectiveness . Key Responsibiliti esDesign, implement, and maintain enterprise observability solutions using Splunk Enterprise including dashboards, alerts, and data ingestion pipelin esDevelop and enhance monitoring frameworks for infrastructure, applications, and web platfor msAutomate operational processes using Linux shell scripting and Pyth onImplement intelligent alerting strategies to reduce noise and improve incident response efficien cyProvide L3 production support for business-critical applications and infrastructu reSupport cloud and containerized deployments across AWS and Kubernetes environmen tsCollaborate with engineering teams to standardize logging and telemetry practic esDrive root cause analysis, post-incident reviews, and continuous reliability improvemen tsBuild operational runbooks, disaster recovery procedures, and service continuity pla nsIntegrate monitoring and deployment workflows with CI/CD tools such as Jenkins, Git, and TeamCi tySupport database monitoring and performance analysis across SQL Server, Oracle, DB2, and MySQL platfor msParticipate in ITIL-based change, incident, and problem management process es Required Ski llsStrong hands-on expertise in Splunk engineering, administration, and architect ureAdvanced experience in Linux / Unix environme ntsProficiency in Python, Shell scripting, and automation framewo rksExperience with AWS cloud services and Kubernetes / Docker platfo rmsKnowledge of monitoring tools such as Nagios and custom observability soluti onsExperience supporting high-availability web platforms and distributed syst emsStrong troubleshooting and production incident management ski llsUnderstanding of CI/CD pipelines and deployment automat ionFamiliarity with ITIL processes and service management tools like Service Now Preferred Qualificat ionsSplunk certifications (Power User / Admin / Archit ect)Experience building large-scale telemetry platf ormsBackground in financial services or high-transaction enterprise environm entsExperience designing intelligent alerting and automated incident workf l ows Experience L evel15+ years in production engineering / SRE / observability r olesPrior experience supporting mission-critical enterprise sys tems
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago