
Remote Senior Site Reliability Engineer
RoShay Services · Jersey City, NJ
Remote
About the job
We are seeking a Senior Site Reliability Engineer to enhance the reliability, scalability, and observability of cloud-based systems.
Key Responsibilities
Improve service delivery and system reliability throughout the entire lifecycle
Monitor and measure system health, availability, and latency
Identify and resolve errors and instability in production cloud services
Collaborate with product and platform teams to enhance system resilience and observability
Innovate and automate to reduce operational toil
Participate in on-call duties as required
Qualifications & Skills
Experience designing, implementing, and operating observability systems in cloud environments
Proficiency with Configuration Management and Infrastructure as Code tools like Terraform or Ansible
Knowledge of cloud platforms (AWS, Azure), containerization, and orchestration technologies
Experience with APM and observability tools such as New Relic, Splunk, Prometheus, Grafana
Background in Linux Systems Engineering and enterprise continuous delivery environments
Development skills in JavaScript, Node.js, or TypeScript
Familiarity with incident response tools and practices in a blameless environment
Strong understanding of security best practices and cloud design patterns for scalability and resiliency
Ability to work autonomously within a distributed team
This role offers a competitive salary, comprehensive benefits, and a remote-first environment focused on collaboration, growth, and innovation.
Ready to apply?Apply now