JJobsSonar

Senior DevOps / SRE (Platform Reliability Engineer) - French fluent

emagine · Portugal

DevOpsRemote

About this role

The Senior DevOps / SRE will ensure the reliability, scalability, and security of the platform and cloud infrastructure. Responsibilities include designing and maintaining scalable AWS infrastructure, improving system reliability through SRE practices, building CI/CD pipelines, managing container orchestration platforms, and implementing monitoring and security solutions. The role also involves incident response, cost optimization, and collaboration with development teams to enhance deployment strategies and system resilience.

Skills & technologies

Must have

  • AWS
  • Kubernetes
  • Docker
  • Helm
  • Terraform
  • Ansible
  • CloudFormation
  • CI/CD
  • GitLab CI
  • Jenkins
  • GitHub Actions
  • Azure DevOps
  • Prometheus
  • Grafana
  • ELK
  • Datadog
  • Splunk
  • Bash
  • Python
  • Linux

Nice to have

  • Azure
  • GCP

Read full description

About the job We are looking for a Senior DevOps / Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and security of our platform and cloud infrastructure. You will play a key role in building and operating cloud-native systems, improving observability, automating operations, implementing SRE best practices (SLOs/SLIs), and supporting development teams to deliver highly available services. Key Responsibilities Design, implement, and maintain highly available and scalable infrastructure on AWS. Own and improve the reliability of production systems using SRE principles (SLO, SLI, error budgets). Build and manage CI/CD pipelines to support fast and safe software delivery. Develop and maintain Infrastructure as Code (IaC) using Terraform, Ansible, CloudFormation, etc. Manage and optimize container orchestration platforms (Kubernetes, Docker, Helm). Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK, Datadog, Splunk). Lead incident response, perform root cause analysis, and write postmortems to drive continuous improvement. Improve system performance, capacity planning, scaling strategies, and disaster recovery processes. Collaborate closely with development teams to improve deployment strategies and system resilience. Implement security best practices (IAM, secret management, vulnerability scanning, patching). Define operational standards, runbooks, documentation, and best practices for platform reliability. Participate in on-call rotation and provide senior-level support for critical production issues. Key Responsibilities (5 Main Missions) The DevOps / SRE lead will be responsible for the stability and evolution of the platform. Your role is structured around five main areas: Mission 1: AWS Infrastructure Management (Build & Run) Mission 2: CI/CD and Deployment Automation Mission 3: Monitoring, Observability, and Alerting: Global Monitoring, Log Management, Application Monitoring, Business Analytics Mission 4: Incident Management, Resilience, and Security Mission 5: FinOps and AWS Cost Optimization Key Requirements 5+ years of experience in DevOps / SRE / Cloud Infrastructure / Platform Engineering. Strong expertise in Linux systems administration and troubleshooting. Proven experience with Kubernetes in production environments. Strong experience with CI/CD tools (GitLab CI, Jenkins, GitHub Actions, Azure DevOps). Solid knowledge of Infrastructure as Code (Terraform highly preferred). Experience with AWS cloud platforms. Strong understanding of networking fundamentals (TCP/IP, DNS, load balancing, reverse proxies). Experience with observability tools: monitoring, metrics, logging, tracing. Strong scripting skills (Bash, Python, or similar). French advanced level. Nice to Have Experience with additional cloud platforms (Azure, GCP). Strong understanding of networking fundamentals.
Ready to apply?Apply now

Similar DevOps jobs

All DevOps jobs