JJobsSonar

Observability & SRE Engineer

Ubique Systems · Germany

SRE / ReliabilityRemote

About this role

The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. The role involves deploying and maintaining Prometheus, Grafana, Loki, and Jaeger stacks as code. Key responsibilities include developing PromQL/LogQL monitoring dashboards and configuring AlertManager routing for autonomous Air-Gap operations. The engineer will also implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation. Additionally, the position requires configuring secure log export and security event management for external SIEM platforms.

Skills & technologies

Must have

Read full description

Role : Observability & SRE Engineer


Mode: 1 year Fixed term contract

We need german speaking candidates min at C1 level

Remote but candidate based in Germany are preferred


Role: Observability & SRE Engineer

Required Clearnce (need to clear before joining): Public Sector Clearance + NdK (Platform Telemetry Vetting)


Primary Skills: Prometheus, Grafana, Loki, OpenTelemetry SDKs, Jaeger tracing, SIEM log export.


  • Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)

Position Overview

The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations.


Key Responsibilities

  • Monitoring Stack Setup: Deploy and maintain Prometheus, Grafana, Loki, and Jaeger stacks as code.
  • Dashboards & Alerting: Develop PromQL/LogQL monitoring dashboards; configure AlertManager routing and inhibition for autonomous Air-Gap operations.
  • Distributed Tracing: Implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation.
  • SIEM Integration: Configure secure log export and security event management (CEF/Syslog) for external SIEM platforms.

Technical Qualifications & Skills

  • Must Have:
  • Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.
  • Hands-on experience with Loki log aggregation and AlertManager routing.
  • Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.
  • Knowledge of SIEM integrations and security event logging.
  • Scripting skills in Go, Python, Shell, and YAML.


Ready to apply?Apply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs