JJobsSonar

Site Reliability Engineer

Selby Jennings · Zurich, Zurich, Switzerland

SRE / ReliabilityOn-site

About this role

The Site Reliability Engineer will be responsible for ensuring the availability, stability, and performance of Linux-based trading systems in a low-latency, high-performance environment. This role involves incident management, on-call duties, and leading post-mortems to drive automation and prevent recurrence. The individual will partner with developers and traders to ensure reliable, high-performance system design and deployment. Responsibilities include low-level system tuning, performance diagnostics, and delivering infrastructure as code using tools like Ansible, Terraform, and Python. The role requires deep expertise in Linux internals, networking, and automation.

Skills & technologies

Must have

  • Linux
  • Ansible
  • Terraform
  • Python
  • Docker
  • Prometheus
  • Grafana
  • ELK
  • perf
  • ftrace
  • tcpdump
  • eBPF
  • YAML
  • JSON
  • Git

Read full description

About the job Our client, a leading proprietary trading firm specialising in both systematic and discretionary strategies, is seeking a Site Reliability Engineer to join their Zurich office. This is a unique opportunity to evolve and enhance a highly sophisticated production trading environment, ensuring exceptional uptime and performance. The role focuses on delivering code-driven solutions while partnering closely with developers and traders to strengthen reliability, observability, and overall operational maturity within a low-latency, high-performance ecosystem. The ideal candidate will bring deep experience supporting highly available, performance-critical, latency-sensitive systems, alongside a strong understanding of Linux internals and networking. A solid background in reliability engineering is essential, with a clear automation-first mindset and hands-on experience with containerisation technologies. Key responsibilities: * Reliability & Production Ownership: Own availability, stability, and performance of Linux-based trading systems (RedHat, Rocky, Ubuntu). * Incident Response: Lead incident management, on-call, and blameless post-mortems, driving automation to prevent recurrence. * Operational Processes: Maintain runbooks, documentation, and standards for consistent production support. * Production Readiness: Partner with developers and traders to ensure reliable, high-performance system design and deployment. * Linux Systems & Performance: Perform low-level tuning (CPU, IRQ, memory, networking) for latency-sensitive workloads. * Performance Diagnostics: Troubleshoot using perf, ftrace, tcpdump, and eBPF. * Automation & Infrastructure: Deliver infrastructure as code with Ansible, Terraform, Python, and shell scripting. Required Qualifications: * Experience in Site Reliability Engineering, Linux engineering, DevOps, or infrastructure-focused roles. * Production Systems: Proven experience supporting highly available, performance-sensitive production environments. * Linux Expertise: Deep knowledge of Linux internals, including scheduling, memory management, interrupts, filesystems, and storage. * Networking: Strong understanding of TCP/IP, UDP, multicast, and distributed systems networking. * Automation & Tooling: Proficiency with Ansible, Terraform, Python, shell scripting, YAML/JSON, and Git-based workflows. * Containers & Observability: Experience with Docker (or similar) and familiarity with observability tools such as Prometheus, Grafana, ELK, or equivalent.
Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs

Site Reliability Engineer Sr

Dayforce · United States

SRE / ReliabilityRemoteEasy apply$80.5K/yr - $143.8K/yr2w ago

Senior SRE (Cloud)

Hazelcast · United Kingdom

SRE / ReliabilityRemoteEasy apply1mo ago