
DevOps Engineer (GCP)
Satori Analytics · Greece
DevOpsRemote
About this role
The DevOps / Platform Engineer will own and evolve the infrastructure for an AI agent evaluation platform, ensuring reliability, observability, security, and fast deployment. They will work closely with backend, ML, and frontend engineers to streamline operations and improve system resilience.
Skills & technologies
Must have
- Terraform
- GCP
- Kubernetes
- Docker
- GitHub Actions
- Python
- PostgreSQL
- Redis
- Celery
- Grafana
- ClickHouse
- LiteLLM
Nice to have
- Celery
- Redis
- LLM/AI infrastructure
- Observability tooling
- Security/compliance
- Cost-optimisation
Read full description
About the job
Are you passionate about AI? 🤖
At Satori Analytics, we aim to change the world one algorithm at a time by bringing clarity to global brands through Data & AI. From cloud-based ecosystems for fintech to predictive models for airlines, our cutting-edge solutions cover the entire data lifecycle—from ingestion to AI applications.
As a fast-growing scale-up, our team of 100+ tech specialists—including Data Engineers, Data Scientists, and more—delivers innovative analytics solutions across industries like FMCG, retail, manufacturing and FSI. Join us as we lead the data revolution in South-Eastern Europe and beyond!
Together with a partnering company, we're looking for a a DevOps / Platform Engineer to own and evolve the infrastructure that keeps this platform reliable (AI agent evaluation platform), observable, secure, and fast to ship to. You'll work closely with backend, ML, and frontend engineers to make deploying and operating services boring, repeatable, and safe.
What Your Day Might Look Like:
Cloud infrastructure as code: Own and extend our Terraform estate across multiple GCP environments (base, core, obs, dev, test, prod), including GKE clusters, Cloud SQL (Postgres/MySQL), networking, buckets, and IAM. Drive the in-progress "Neo" platform rollout and the cutover/retirement of legacy infrastructure
Kubernetes & containers: Manage workloads on GKE, maintain Dockerfiles and Helm-style application configs for :10 backend services, and tune autoscaling, resource limits, and pod disruption budgets
Maintain and improve our GitHub Actions pipelines: PR checks (Python/JS lint, type-check, tests), Terraform prechecks, image builds and pushes, auto-deploy, and DB-migration labelling/gating. Reduce build times and flakiness, and make deploys self-service for product teams
Data & messaging infrastructure: Operate Postgres, Redis, and Celery-based async workers; manage Alembic migrations, queue health, and backpressure for long-running simulation jobs
Observability: Own our monitoring stack — Grafana dashboards, ClickHouse, Langfuse (LLM tracing), and Celery queue metrics. Build alerting and SLOs so we catch issues before customers do
Security & secrets: Manage secret distribution, least-privilege IAM, and remediation tracking. Partner with engineering on findings in our security assessment process
Cost & reliability: Keep an eye on cloud and LLM-proxy (LiteLLM) spend, right-size resources, and improve resilience of the simulation and evaluation pipelines
You'll work with:
Cloud: Google Cloud Platform (GKE, Cloud SQL, GCS, IAM); some AWS / IBM footprint
IaC: Terraform (>= 1.14), multi-environment root modules
Containers/orchestration: Docker, docker compose (local), Kubernetes / GKE
CI/CD: GitHub Actions
Backend: Python 3.13+ (managed with uv), Celery, FastAPI-style HTTP APIs; Node/Express services
Data: PostgreSQL, MySQL, Redis, ClickHouse
Observability: Grafana, Langfuse, custom Celery metrics
LLM infra: LiteLLM proxy
Requirements
Your Superpowers 🚀
3+ years in DevOps / SRE / Platform Engineering, or strong backend experience with heavy infra ownership
Solid hands-on Terraform (modules, state, multi-environment) and cloud experience (GCP preferred; AWS/Azure transferable)
Production Kubernetes experience: deployments, services, autoscaling, debugging pods, rollouts/rollbacks
Strong Docker fundamentals and comfort writing/optimising Dockerfiles
CI/CD pipeline design and maintenance (GitHub Actions, or equivalent like GitLab CI / CircleCI)
Comfortable scripting and reading code in Python and/or Bash; able to navigate a polyglot monorepo
Operational experience with relational databases and managed database services (migrations, backups, performance)
A reliability mindset: monitoring, alerting, incident response, and writing runbooks
Bonus points for:
Experience operating Celery / distributed task queues and Redis at scale
Familiarity with LLM/AI infrastructure (model proxies, GPU scheduling, token/cost management)
Observability tooling depth (Grafana, Prometheus, ClickHouse, OpenTelemetry, Langfuse or similar tracing)
Security/compliance experience (IAM hardening, secret management, vulnerability remediation)
Cost-optimisation experience for cloud + third-party API spend
Experience supporting a monorepo with multiple language ecosystems and editable/internal package dependencies
Benefits
Perks on Perks
Competitive salary
Training budget to level up your skills from top tech partners like Microsoft, AWS, Salesforce, and Databricks - whether it's certifications or courses, we've got you covered
Private insurance, top-tier tech gear, and the chance to work with a stellar crew
Ready to create some data magic with us? Hit that apply button and let's get started. ✨
Ready to apply?Apply now