
Senior Site Reliability Engineer
Ladders · United States
Remote
About the job
For our client, we are seeking a Senior Site Reliability Engineer to join the team of a leader in the Transportation space. This role will lead technical work focused on cloud-enabled scalability, reliability, and delivery excellence. You will work across engineering, product, operations, and business stakeholders to translate complex requirements into practical technology solutions. The position offers the opportunity to influence architecture, execution quality, and the technology capabilities that enable long-term growth within a operationally complex logistics environment.
Location: Remote - US based candidates only, no visa sponsorship available
Compensation: $120,000 – $150,000 annually
Responsibilities
Own architecture and implementation of infrastructure on GCP, ensuring reliability and scalability
Manage containerized workloads on Kubernetes, focusing on performance and resource optimization
Automate processes to reduce manual toil and improve efficiency
Develop and maintain observability solutions with monitoring and alerting mechanisms
Design CI/CD pipelines to ensure safe and efficient deployments
Collaborate with teams to enhance deployment practices and ensure application reliability
Participate in incident response, leading post-mortem reviews and implementing durable fixes
Qualifications
5+ years in SRE, platform, or infrastructure engineering with ownership of complex systems
Strong hands-on experience with GCP core services: GKE, Cloud Run, AlloyDB, and networking
Proficient in Docker and Kubernetes for troubleshooting and scaling workloads
Deep experience with Terraform for Infrastructure as Code, writing reusable modules
Ability to program in TypeScript, Python, Go, or similar for automation tooling
Strong experience with action-oriented monitoring and alerting using tools like Datadog or Prometheus
Solid grasp of distributed systems fundamentals including failure modes and consistency
Benefits
High-impact infrastructure ownership and visible contributions in a small team setting
Trust in your judgment with a direct path from decision-making to production
Opportunities to shape reliability and automation practices during a growth phase
A collaborative environment emphasizing clear communication and shared ownership
Our client is an equal opportunity employer. We encourage you to apply even if you don’t meet every qualification—your background could be exactly what this team needs.
Desired Skills and Experience
Google Cloud Platform (GCP), GKE (Google Kubernetes Engine), Cloud Run, AlloyDB, Terraform, Docker, Kubernetes, TypeScript, Python, Go, Datadog, Git, PostgreSQL, Redis, ClickHouse, Kafka, Redpanda, Event-driven architecture tools (Pub/Sub), CI/CD systems (GitHub Actions), Observability tools (Prometheus, Grafana), Infrastructure as Code best practices, Cloud cost management tools, Disaster recovery techniques, Business continuity planning, Production incident management processes
$120K/yr - $150K/yrApply now