JJobsSonar

Site Reliability Engineer/GCP

Filevine · United States

SRE / ReliabilityRemote

About this role

The Site Reliability Engineer will lead reliability engineering efforts, design and maintain autonomous systems for software development and operations, and ensure system availability, performance, and security. They will also mentor junior engineers, participate in on-call rotations, and drive continuous improvement in CI/CD pipelines and reliability practices.

Skills & technologies

Must have

  • Python
  • Bash
  • PowerShell
  • Google Cloud Platform (GCP)
  • Compute Engine
  • Kubernetes Engine/GKE
  • Cloud Monitoring
  • Cloud Logging
  • IAM
  • CI/CD
  • Automation
  • Monitoring/Alerting
  • Incident Response
  • Capacity Planning
  • Performance Optimization

Nice to have

  • M.S. in computer science
  • Google Cloud Professional certifications
  • AWS certifications
  • Artificial Intelligence
  • Machine Learning
  • Internet scale applications

Read full description

About the job Filevine is a fast-growing SaaS startup in need of a Site Reliability Engineer to be part of our growing development team. Filevine is a perfect stage to join because you can be a major player driving the direction and strategy of our growth. You will apply your knowledge of software architecture and development with a team of engineers building large features and innovation according to provided design specifications. Responsibilitie sProvide strong leadership, mentoring, and sound judgment as the Reliability Engineering lead on your team .Design and maintain autonomous systems for building, deploying, testing, and operating all Filevine products .Act as the authoritative voice of reliability across the full software development lifecycle (SDLC) .Monitor, aggregate, dashboard, and alert on software/infrastructure events to ensure visibility and fast response .Continuously enhance CI/CD pipelines, automation scripts, playbooks, and tools to streamline processes and reduce resolution time .Proactively identify and resolve gaps in system availability, performance, and security while defending overall security posture .Document processes, architecture, procedures, and best practices; research, adopt, or build reliable tools to boost engineer productivity .Collaborate within your team (or independently), mentor junior engineers, participate in 24/7 on-call rotation for production support and emergency response, and communicate clearly with technical and management stakeholders .Qualification s8+ years of hands-on technical experience in software engineering, infrastructure, or operations roles, including a minimum of 6 years dedicated to Site Reliability Engineering (SRE) .Demonstrated curiosity, self-motivation, continuous learning mindset, passion for improvement, and proactive enthusiasm to enhance systems and processes daily without needing direction .Strong proficiency in Python, Bash, PowerShell, and other common SRE tooling and scripting technologies .Expert-level experience designing, building, and maintaining autonomous systems that handle software build, deployment, testing, monitoring, and operations with minimal human intervention .Deep proficiency with Google Cloud Platform (GCP) and its core SRE services, including Compute Engine, Kubernetes Engine/GKE, Cloud Monitoring, Cloud Logging, and IAM. Experience with AWS is a strong plus (e.g., EC2, EKS, CloudWatch, S3) .Proficiency in all core skills expected of an SRE II, including monitoring/alerting, incident response, capacity planning, performance optimization, CI/CD pipeline enhancement, and reliability engineering best practices .Bachelor’s degree in Computer Science, Information Systems, or a related field; equivalent certifications (e.g., Google Cloud Professional certifications, AWS certifications); or substantial comparable direct work experience .Proven track record of independently driving reliability improvements, reducing toil through automation, and contributing to high-availability, scalable production systems in a fast-paced environmen tNice-to-Hav eM.S. in computer science, information systems, a related field; comparable certifications; or equivalent direct work experienc eExperience developing, deploying, and maintaining internet scale application sExperience incorporating Artificial Intelligence or Machine Learning into internet scale application s
Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago

Senior SRE (Cloud)

Hazelcast · United Kingdom

SRE / ReliabilityRemoteEasy apply1mo ago