JJobsSonar

AI-DNA Senior Infrastructure & Reliability Engineer

IgniteTech · Germany

SRE / ReliabilityRemote

About this role

The Senior Infrastructure & Reliability Engineer will own production reliability and build and improve AI agents for incident triage, triage, and auto-remediation. The role involves executing production deployments and infrastructure changes with strong operational discipline. The engineer will write production code, automation, and AI workflows to eliminate repetitive operational work. The role is part of an continuous improvement of an AI-native Infrastructure & Reliability organization. The platform is-built on AWS and-on an enterprise-scale SaaS platform. The engineer will work within a automation-first culture and-on an AI-first culture.

Skills & technologies

Must have

  • AWS
  • Infrastructure Engineering
  • actionable-automation
  • Incident Response
  • Claude Code
  • Codex
  • engineering-tools
  • Cloud Operations
  • Platform Engineering

Read full description

We're building an AI-native Infrastructure & Reliability organization where autonomous agents investigate incidents, validate hypotheses, generate RCAs, and safely remediate production issues.

We're looking for a Senior Infrastructure & Reliability Engineer to help build and operate this platform. You'll own production reliability while continuously improving the AI agents that automate incident response, operational workflows, and infrastructure management.

What You'll Do

  • Own production reliability, responding to and resolving customer-impacting incidents.
  • Build and improve AI agents for incident triage, change validation, RCA generation, and auto-remediation.
  • Execute production deployments and infrastructure changes with strong operational discipline.
  • Write production code, automation, runbooks, and AI workflows that eliminate repetitive operational work.
  • Continuously improve platform uptime, operational efficiency, and customer experience.

What We're Looking For

  • 5+ years operating enterprise SaaS platforms in Infrastructure, Platform Engineering, DevOps, Cloud Operations, or Site Reliability Engineering.
  • Strong AWS production experience, including large-scale cloud infrastructure and incident response.
  • Hands-on engineer who enjoys troubleshooting complex production systems and writing automation.
  • Experience using AI engineering tools such as Claude Code, Codex, Cursor, Warp, or similar.
  • Passion for building AI-powered operations, automation, and autonomous infrastructure.
  • Excellent written and spoken English.

Why Join Us

You'll help build one of the industry's first AI-native Infrastructure & Reliability organizations, where engineers are measured not by the number of tickets they close, but by the AI systems they build to eliminate them.

  • AI-first culture
  • No limits on AI tooling or compute
  • Fully remote
  • Global team
  • Enterprise-scale SaaS
$100,000.00/yr - $100,000.00/yrApply now

Similar SRE / Reliability jobs

All SRE / Reliability jobs