JJobsSonar

DevOps / Platform Engineer

MincaAI · Argentina

Remote
About the job MincaAI — Remote, LatAm-based, freelance About MincaAI MincaAI is an AI-native insurance software company. HQ at Station F (Paris) x Mexico, operations across LatAm. We build production AI systems for enterprise insurance carriers across three pillars — underwriting, policy issuance, and billing automation. Seven microservices in production. Real LLM pipelines. Real enterprise clients. The Role We ship the same product into two very different environments, and this is the heart of the role: Internal: a self-managed Kubernetes cluster on AWS, delivered with Helm and GitOps (ArgoCD). Client-hosted production: a managed Red Hat OpenShift cluster, delivered with OpenShift-native primitives. You own reliability, security, and delivery velocity across both. From commit to running observable workload — no handoffs, no ambiguity. This is a builder role with real ownership. You push back on complexity, retire tech debt, and set the standards other engineers deploy against. Architecture is not one-size-fits-all. We operate a SILO model — each client engagement has a different infra topology. With some clients (e.g. GdS), we coordinate closely with their in-house infrastructure team. With others, we own STG and Prod end-to-end and get to be creative. You need to be comfortable pivoting between these modes and pragmatic about which pattern fits which client. There is no template we deploy blindly. Workload split. This engagement is roughly 80% platform / DevOps (K8s, OpenShift, GitOps, CI/CD, observability) and 20% code quality ownership via SonarQube — managing reports across services, enforcing quality gates in CI, driving remediation with dev teams, and tracking tech-debt reduction. Both halves matter equally; a strong DevOps profile who won't touch quality gates is not the right fit. What You'll Own Operate and evolve the self-managed K8s cluster on AWS (ingress, TLS automation, bootstrap tooling). GitOps with ArgoCD: sync policies, drift/self-heal, automated image reconciliation. Helm charts: zero-downtime rollouts, HPA, PDBs, quotas, default-deny network policies, migration hooks. OpenShift deployments client-side: build system, internal registry, routes, SCCs, arbitrary-UID images. CI/CD on GitHub Actions with self-hosted runners: path-aware pipelines, semantic releases, environment gates. Secrets & supply chain: SOPS + AWS KMS, IAM hygiene, secret scanning, immutable tags, fail-closed defaults. AWS + Terraform: managed Postgres/Redis, object storage with tenant isolation, KMS, ECR, pod-level IAM. Observability: OpenTelemetry (logs/traces/metrics), Prometheus-compatible backend, alerting on AI workloads. Data plane support: Postgres + pgvector, Redis streams, gated migrations. SonarQube ownership (~50% of the engagement): manage SonarQube reports across all services, define and enforce quality gates in CI, triage and prioritize findings (bugs, vulnerabilities, code smells, coverage), and drive remediation workflows with the dev teams. Turn static analysis into shipped improvements, not shelf-ware. DevEx & CLI tooling: contribute to internal developer tooling, CLIs, and workflows that make the whole team faster. Continuous improvement, not one-off setup: ongoing maintenance and evolution of infrastructure and workflows as the product and client base grow. Cross-team collaboration: work with engineers across product teams on architecture proposals and design reviews from the DevOps / infra side. You review, you get reviewed, you don't operate solo. Runbooks, SLOs, incident response, DR. Must-Have Production Kubernetes on a self-managed cluster (not just EKS/GKE UX). Helm charts written and maintained end-to-end. GitOps with ArgoCD (or Flux). GitHub Actions at scale, including self-hosted runners. AWS + Terraform with remote state and IAM hygiene. Secrets management with SOPS+KMS (or equivalent). Docker multi-stage, digest pinning, registry workflows. Bash + enough Python to debug a FastAPI microservices backend. OpenTelemetry and a real observability stack in production. Uses AI coding tools (Cursor, Claude Code, or equivalent) as part of your daily workflow. Hands-on SonarQube experience: configuration, quality gates in CI, report triage, and driving remediation with dev teams (not just producing dashboards). Architectural flexibility: comfortable adapting across different client topologies — from full STG/Prod ownership to coordinating with a client-side infra team. Pragmatic and creative, not dogmatic. Nice-to-Have OpenShift specifics (oc, build configs, image streams, SCCs). Postgres tuning + pgvector, Redis Streams consumer groups. SBOM / image scanning / supply-chain security. Multi-tenant SaaS with staging/prod parity across self-hosted + client-managed clusters. Observability of LLM workloads (cost, latency, drift). Our Stack Orchestration: Kubernetes (self-managed), OpenShift, Helm, ArgoCD CI/CD: GitHub Actions on self-hosted runners Cloud/IaC: AWS (ECR, S3, KMS, RDS, ElastiCache, IAM), Terraform Secrets: SOPS + AWS KMS Data: PostgreSQL + pgvector, Redis (streams + cache), S3 Observability: OpenTelemetry, Grafana, Prometheus-compatible store Apps: Python / FastAPI microservices, Vite + React + TypeScript How We Work SOLID, DRY, KISS. Small focused changes. No new tech debt. If a request conflicts with best practice, you push back and propose better. Reproducible by default: IaC + GitOps as source of truth. Reproduce before you fix: bugs get captured in a test before the patch lands. Success Deploys on both clusters are one-click, observable, reversible. Secrets, image, and IAM hygiene pass audit with fail-closed defaults. CI is fast and trustworthy — flakiness and manual steps are gone. Autoscaling and quotas tuned. Environments stable under load, no over-provisioning. Runbooks and SLOs exist. On-call is calm. Logistics Contract: Freelance, 4 months (extension possible based on fit and roadmap). Location: Remote, LatAm-based (Argentina / Colombia / Mexico preferred). Languages: Professional English required. Spanish required. Timezone: Overlap with LatAm working hours. Start: ASAP. How to Apply Send your CV + a short note (5 lines max) on the hardest production incident you personally led to resolution, to stella.choi@mincaai.com
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago