
Site Reliability Engineer
Data Nexus AI · United States
Remote
About the job
This is a W2 Role, Please do not apply for C2C
We are looking for a highly experienced Senior Observability & Site Reliability Engineer to support large-scale enterprise platforms and mission-critical applications. The ideal candidate will have deep hands-on experience in building and operating end-to-end monitoring, logging, and alerting solutions across distributed environments
.This role involves close collaboration with development, infrastructure, and operations teams to ensure platform reliability, performance visibility, and incident response effectiveness
.
Key Responsibiliti
esDesign, implement, and maintain enterprise observability solutions using Splunk Enterprise including dashboards, alerts, and data ingestion pipelin
esDevelop and enhance monitoring frameworks for infrastructure, applications, and web platfor
msAutomate operational processes using Linux shell scripting and Pyth
onImplement intelligent alerting strategies to reduce noise and improve incident response efficien
cyProvide L3 production support for business-critical applications and infrastructu
reSupport cloud and containerized deployments across AWS and Kubernetes environmen
tsCollaborate with engineering teams to standardize logging and telemetry practic
esDrive root cause analysis, post-incident reviews, and continuous reliability improvemen
tsBuild operational runbooks, disaster recovery procedures, and service continuity pla
nsIntegrate monitoring and deployment workflows with CI/CD tools such as Jenkins, Git, and TeamCi
tySupport database monitoring and performance analysis across SQL Server, Oracle, DB2, and MySQL platfor
msParticipate in ITIL-based change, incident, and problem management process
es
Required Ski
llsStrong hands-on expertise in Splunk engineering, administration, and architect
ureAdvanced experience in Linux / Unix environme
ntsProficiency in Python, Shell scripting, and automation framewo
rksExperience with AWS cloud services and Kubernetes / Docker platfo
rmsKnowledge of monitoring tools such as Nagios and custom observability soluti
onsExperience supporting high-availability web platforms and distributed syst
emsStrong troubleshooting and production incident management ski
llsUnderstanding of CI/CD pipelines and deployment automat
ionFamiliarity with ITIL processes and service management tools like Service
Now
Preferred Qualificat
ionsSplunk certifications (Power User / Admin / Archit
ect)Experience building large-scale telemetry platf
ormsBackground in financial services or high-transaction enterprise environm
entsExperience designing intelligent alerting and automated incident workf
l
ows Experience L
evel15+ years in production engineering / SRE / observability r
olesPrior experience supporting mission-critical enterprise sys
tems
Ready to apply?Apply now