
About this role
The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. The role involves deploying and maintaining Prometheus, Grafana, Loki, and Jaeger stacks as code. Key responsibilities include developing PromQL/LogQL monitoring dashboards and configuring AlertManager routing for autonomous Air-Gap operations. The engineer will also implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation. Additionally, the position requires configuring secure log export and security event management for external SIEM platforms.
Skills & technologies
Must have
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- Jaeger
- SIEM
- Go
- Python
- Shell
- YAML
Read full description
Role : Observability & SRE Engineer
Mode: 1 year Fixed term contract
We need german speaking candidates min at C1 level
Remote but candidate based in Germany are preferred
Role: Observability & SRE Engineer
Required Clearnce (need to clear before joining): Public Sector Clearance + NdK (Platform Telemetry Vetting)
Primary Skills: Prometheus, Grafana, Loki, OpenTelemetry SDKs, Jaeger tracing, SIEM log export.
- Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)
Position Overview
The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations.
Key Responsibilities
- Monitoring Stack Setup: Deploy and maintain Prometheus, Grafana, Loki, and Jaeger stacks as code.
- Dashboards & Alerting: Develop PromQL/LogQL monitoring dashboards; configure AlertManager routing and inhibition for autonomous Air-Gap operations.
- Distributed Tracing: Implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation.
- SIEM Integration: Configure secure log export and security event management (CEF/Syslog) for external SIEM platforms.
Technical Qualifications & Skills
- Must Have:
- Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.
- Hands-on experience with Loki log aggregation and AlertManager routing.
- Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.
- Knowledge of SIEM integrations and security event logging.
- Scripting skills in Go, Python, Shell, and YAML.