
Platform Engineer (Cloud Infrastructure / DevOps / SRE)
AIUTA · Cyprus
Platform EngineeringRemote
About this role
The Platform Engineer will build and scale the infrastructure behind AIUTA’s AI-powered solutions, working across cloud infrastructure, Kubernetes, CI/CD systems, observability, and GPU-powered AI workloads. The role involves designing scalable cloud infrastructure, optimizing CI/CD pipelines, supporting AI/ML workloads, and improving system resilience and cost efficiency.
Skills & technologies
Must have
- Terraform
- Kubernetes
- GitLab CI
- GitHub Actions
- Prometheus
- Grafana
- OpenSearch
- Linux
- Networking
- Security
- CI/CD
- Cloud (AWS/GCP)
- DevSecOps
- GitOps
Nice to have
- Nvidia/AMD GPU
- MLOps
- Ansible
- PagerDuty
- Disaster Recovery
Read full description
About the job
About The Company
AIUTA is a B2B fashion AI infrastructure company that enables brands and retailers to create and scale high-quality visual content and personalized shopping experiences. We power solutions such as virtual try-on, AI Studio, and outfit recommendations, allowing brands to empower their customers to see how products look on themselves or on models similar to them before buying, and to discover, mix and match, and visualize complete outfits. With proven results across leading fashion brands, we combine proprietary AI with human quality control to deliver production-ready outputs at scale, enabling brands to create more, convert better, and operate more efficiently.
About The Role
We’re looking for a Platform Engineer to help us build and scale the infrastructure behind AIUTA’s AI-powered solutions trusted by leading retail players.
You’ll work across modern cloud infrastructure, Kubernetes platforms, CI/CD systems, observability, and GPU-powered AI workloads — partnering closely with engineering and data science teams to improve scalability, reliability, and developer velocity.
This is a highly hands-on role with strong ownership and direct impact on how our platform evolves as we continue scaling our products, infrastructure, and engineering organization.
Key Responsibilities
Architect & Automate: Design, build, and maintain highly scalable, secure, and resilient cloud infrastructure using Infrastructure as Code (IaC) with Terraform.
Streamline Delivery: Build and optimize blazing-fast CI/CD pipelines (GitLab CI, GitHub Actions) to elevate developer experience and deploy containerized services.
Scale AI Workloads: Partner with our AI/ML teams to support high-performance computing infrastructure, optimizing containerized environments (Kubernetes) for GPU/accelerated workloads.
Boost Resilience: Implement modern observability, monitoring, and proactive logging systems (Prometheus, Grafana, OpenSearch/ELK) to maintain high service availability.
Control Costs & Performance: Proactively identify performance bottlenecks and execute cloud cost-optimization strategies without compromising on speed or reliability.
Drive Engineering Culture: Reduce environment drift, advocate for GitOps/DevSecOps best practices, participate in chaos/resiliency testing, and mentor team members as we scale the engineering organization.
What We Are Looking For
Experience: 4+ years of hands-on experience in DevOps, Platform Engineering, or Infrastructure roles, ideally within a fast-growing startup or high-load product environment.
Cloud & Containers: Strong expertise in cloud platforms (AWS or GCP preferred) and deep hands-on experience orchestrating production-grade Kubernetes clusters.
Infrastructure as Code: Proven track record of managing complex environments using Terraform.
Systems & Networking: Solid fundamentals in Linux systems administration, networking, security, and modern web application architectures.
Troubleshooting Mindset: Excellent analytical and debugging skills; a proactive problem-solver who takes ownership of production issues.
Collaborative Spirit: Strong communication skills in English, with the ability to collaborate cross-functionally and document infrastructure patterns clearly.
Nice to Haves
Experience scaling GPU/accelerated hardware (Nvidia/AMD) for AI/ML model inference.
Familiarity with MLOps frameworks or deploying high-performance AI serving stacks.
Experience with configuration management tools (e.g., Ansible) and incident management processes (PagerDuty, disaster recovery planning).
Prior experience working in a fast-paced, Series-A stage startup.
What We Offer
An opportunity to work with cutting-edge AI technology alongside a highly experienced team from leading global tech companies,
Competitive salary,
A collaborative and inclusive team culture where every voice is heard,
Remote-friendly environment,
Flexible working hours.
By submitting this application, I confirm that all the information given by me in this application for employment and any additional documents attached hereto are true to the best of my knowledge and that I have not wilfully suppressed any material fact. I confirm I have disclosed if applicable any previous employment with Aiuta. I accept that if any of the information given by me in this application is in any way false or incorrect, my application may be rejected, any offer of employment may be withdrawn or my employment with Aiuta may be terminated summarily or I may be dismissed. By submitting this application, I agree that my personal data will be processed in accordance with Aiuta's Candidate Privacy Notice
Ready to apply?Apply now