
Site Reliability Engineer
NationsBenefits · United States
Remote
About the job
Position: Site Reliability Engineer II (SRE)
Location: Remote (US-based candidates only)
Employment Type: Full-Time
Department: Site Reliability Engineering
Reports To: VP of Site Reliability Engineering
About NationsBenefits
NationsBenefits is one of America's fastest-growing Healthcare FinTech companies, delivering innovative supplemental benefits, flex card solutions, and member engagement platforms for leading managed care organizations.
Our technology helps health plans improve member outcomes, reduce healthcare costs, and address social determinants of health through secure, scalable, and compliance-driven solutions. With teams across the United States, South America, and India, we foster a collaborative culture that encourages innovation, career growth, and continuous learning.
Position Overview
We are seeking a Site Reliability Engineer II (SRE) to join our growing Site Reliability Engineering team.
In this role, you will help ensure the reliability, availability, and performance of our production platforms by monitoring system health, responding to incidents, troubleshooting Kubernetes-based environments, and collaborating closely with Development, DevSecOps, and Engineering teams.
This position is ideal for engineers who enjoy solving production challenges, automating operational tasks, and improving platform reliability in a fast-paced Healthcare FinTech environment.
Key Responsibilities
Production Support & Incident Management
Serve as the first line of response for production incidents.
Monitor, triage, troubleshoot, and resolve production issues.
Perform initial root cause analysis and escalate incidents when appropriate.
Communicate incident updates to stakeholders throughout the resolution process.
Monitoring & Platform Reliability
Monitor infrastructure and application health using Datadog or similar observability platforms.
Optimize monitoring alerts and reduce false positives.
Troubleshoot Kubernetes workloads, including pods, deployments, logs, and rollbacks.
Maintain high platform availability and performance.
Collaboration
Partner with Development, DevSecOps, Infrastructure, and Engineering teams to resolve production issues.
Participate in cross-functional troubleshooting sessions.
Recommend improvements to monitoring, tooling, and operational processes.
Collaborate effectively with global teams across multiple time zones.
Automation & Continuous Improvement
Develop automation scripts and operational tools using:
Python
PowerShell
Bash
C#
Java
Support CI/CD pipeline monitoring and deployment reliability.
Contribute to self-healing and automated recovery solutions.
Documentation & Compliance
Maintain detailed incident documentation and post-mortem reports.
Follow security and compliance standards including:
HIPAA
PCI DSS
SOC 2
ISO 27001
HITRUST
On-Call & Support Rotation
Participate in weekday production support as part of a global follow-the-sun model.
Participate in an on-call rotation for critical production systems as needed.
Required Qualifications
3–5 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering
Hands-on experience with production incident management and troubleshooting
Experience with Datadog or similar monitoring/observability tools
Strong experience supporting Kubernetes and Docker environments
Experience with SQL, MySQL, or NoSQL databases
Familiarity with cloud platforms such as Azure, AWS, or GCP
Experience working in high-availability production environments
Excellent troubleshooting, analytical, and communication skills
Ability to work weekday shifts in a global follow-the-sun support model
Preferred Qualifications
Experience with CI/CD pipelines and deployment automation
Knowledge of Helm Charts
Understanding of ITIL processes and Agile methodologies
Experience with scripting or programming using:
Python
PowerShell
Bash
C#
Java
Familiarity with security and compliance standards within Healthcare or FinTech environments
Why Join NationsBenefits?
Competitive compensation and comprehensive benefits
Unlimited PTO
Fully remote work environment (US-based)
Opportunity to work with modern cloud-native technologies
Collaborative, supportive, and innovation-driven culture
Continuous learning and career advancement opportunities
Meaningful work that directly impacts healthcare technology and millions of members
Ideal Candidate
We're looking for a proactive engineer who thrives in fast-paced production environments, enjoys solving complex infrastructure challenges, and is passionate about improving system reliability through automation, monitoring, and collaboration. If you have strong Kubernetes, monitoring, and cloud operations experience, we'd love to hear from you.
5 benefitsApply now