JJobsSonar

Site Reliability Engineer

Ubique Systems · Germany

SRE / ReliabilityRemote

About this role

The Site Reliability Engineer will be responsible for platform operations, incident management, and system stability. The role involves managing Kubernetes, container orchestration, and CI/CD pipelines using tools like Jenkins and ArgoCD. Key responsibilities include maintaining Prometheus, Thanos, and Grafana for observability, and administering Elasticsearch clusters and Logstash pipelines. The candidate must be prepared for 24x7 on-call rotations, 24x7 operational support, and participating in Major Incident Management. The position is remote from Germany and requires an EU Passport and Ü2 security clearance.

Skills & technologies

Nice to have

Read full description

Job Title: Site Reliability Engineer (24x7 Operational Support)

Job Type : Permanent

Remote from Germany

Must hold EU Passport


About the Role:


We are seeking a Site Reliability Engineer (SRE) with a strong background in observability, secure logging, and automation. The ideal candidate will have hands-on experience with Elasticsearch and/or Prometheus platforms. This role encompasses critical responsibilities in platform operations, including incident management, execution of scheduled maintenance, and contributing to engineering tasks focused on enhancing system stability. The SRE will also be responsible for adhering to standard operating procedures (SOPs) and actively contributing to their continuous improvement by providing constructive feedback.


Key Responsibilities:

Platform Engineering & DevOps: Manage Kubernetes and container orchestration, including Helm chart configurations and CI/CD pipelines (Jenkins, ArgoCD). Develop automation scripts (Python, Bash, Go) and deploy Infrastructure-as-Code (IaC) solutions.

Observability, Monitoring & Visualisation: MaintainPrometheus solutions (scrape configurations, alert rules, PromQL queries), administer Thanos and Grafana.

Elastic Stack Operations & Log Management: Configure and optimise Elasticsearch clusters, Logstash pipelines, and Kibana dashboards for secure, scalable log processing.

Incident Response, Troubleshooting & Collaboration: Participate in 24x7 on-call rotations for rapid incident response, troubleshoot platform, data and performance issues, and engage in Major Incident Management (MIM).

Secure Operations & Compliance: Ensure system operations meet security and data protection requirements, maintainsecure documentation, and manage access control policies.


Qualifications, Requirements, and Skills

Strong grasp of Linux concepts, preferably in Kubernetes environments.

Solid understanding of networking fundamentals and REST APIs.

Proficiency in Python, Go, or Bash.

Proficiency in Git-based configuration management workflows.

Familiarity with CI/CD tools like Helm, Jenkins, or ArgoCD.

Experience with Elasticsearch and/or OpenSearch.

Fluent English communication skills.

Willingness to work shift-based 24x7 on-call support, including weekends and holidays.

Must possess Ü2 security clearance.

AWS, Azure Knowledge

Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs