JJobsSonar

Software Engineer

Ubique Systems · Germany

SRE / ReliabilityRemote

About this role

The Observability & SRE Engineer will build and operate central monitoring and telemetry platforms for distributed cloud environments. The role involves working with a stack including Prometheus, Grafana, Loki, and OpenTelemetry to manage metrics, logs, and distributed tracing. Key responsibilities include creating monitoring dashboards, configuring AlertManager, and implementing OpenTelemetry Collectors and SDKs. The engineer will work closely with SRE, DevOps, Cloud, Security, and Platform teams to improve system reliability and troubleshoot performance issues. The position is a 1-year fixed-term contract.

Skills & technologies

Must have

Read full description

Observability & SRE Engineer

Location: Remote – Germany preferred

Contract: 1-Year Fixed Term Contract

Salary: Up to €80,000 – Flexible for strong candidates


About the Role

We are looking for an experienced Observability & SRE Engineer to build and operate central monitoring and telemetry platforms for distributed cloud environments.

You will work with metrics, logs, distributed tracing and alerting, helping engineering teams understand system health, troubleshoot issues and improve reliability.

The role involves working with Prometheus, Grafana, Loki, OpenTelemetry and Jaeger, including environments where operations need to work independently or in air-gapped environments.


What You'll Do

  • Build and maintain observability and monitoring platforms using Prometheus, Grafana, Loki and Jaeger.
  • Create monitoring dashboards and alerts using PromQL and LogQL.
  • Configure AlertManager for alert routing and notification.
  • Implement OpenTelemetry Collectors and SDKs for application monitoring and distributed tracing.
  • Support OTLP telemetry collection and trace propagation.
  • Implement distributed tracing using Jaeger.
  • Configure secure log forwarding and integration with external SIEM platforms.
  • Work with security event logging using Syslog / CEF.
  • Support monitoring and telemetry solutions for distributed and air-gapped environments.
  • Automate deployments and configuration using Infrastructure-as-Code and scripting.
  • Troubleshoot monitoring, logging, alerting and performance issues.
  • Work closely with SRE, DevOps, Cloud, Security and Platform teams.


Essential Skills

  • Strong hands-on experience with Prometheus.
  • Strong experience with Grafana and dashboard development.
  • Experience with PromQL.
  • Experience with Loki / LogQL or other centralised log aggregation platforms.
  • Strong understanding of OpenTelemetry, including Collectors, SDKs and OTLP.
  • Experience with Jaeger / distributed tracing.
  • Experience with AlertManager and alert routing.
  • Understanding of SIEM integration and security event logging.
  • Experience with Syslog, CEF or similar security logging formats.
  • Strong troubleshooting and monitoring skills.
  • Scripting experience with Go, Python, Shell and/or YAML.

Clearance Requirement

Candidates must be able to obtain/clear:

  • Public Sector Clearance
  • NdK – Nachweis der Kundigkeit


If you're an experienced SRE, Observability Engineer, DevOps Engineer or Platform Engineer with strong Prometheus/Grafana experience, we'd be interested in hearing from you.


Please apply or send your latest CV via LinkedIn message.


#SRE #SiteReliabilityEngineer #Observability #ObservabilityEngineer #DevOps #PlatformEngineer #Prometheus #Grafana #OpenTelemetry #Jaeger #Loki #Kubernetes #CloudNative #Monitoring #DistributedTracing #SIEM #DevOpsJobs #SREJobs #GermanyJobs #RemoteJobs #Germany

€80,000 / yearApply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs