JJobsSonar

MLOps Engineer

Evlo AI · New York, NY

Remote
About the job About The Role The role bridges the gap between machine learning development and production-grade operations, owning the infrastructure, pipelines, and deployment frameworks that power real-time AI models. The engineer will focus on building robust MLOps platforms to automate model training, deployment, monitoring, and scaling at enterprise volume. Working alongside data scientists, data engineers, and backend developers, the position plays a critical role in standardizing the ML lifecycle. The ideal candidate ensures that models run reliably, cost-effectively, and with minimal latency in a production cloud environment. Key Responsibilities Design, build, and maintain robust CI/CD pipelines for machine learning models, enabling seamless and automated transitions from research to production Deploy and orchestrate machine learning workflows using Kubernetes, Kubeflow, or Apache Airflow to manage complex DAGs and training pipelines Implement comprehensive model monitoring, logging, and alerting systems to detect data drift, concept drift, and performance degradation in real-time Optimize model serving infrastructure utilizing frameworks like Triton Inference Server, TorchServe, or TF Serving to minimize latency and infrastructure costs Build and maintain a centralized feature store (e.g., Feast or Tecton) to ensure consistent data definitions across training and real-time serving environments Collaborate with security and compliance teams to enforce data governance, model lineage, and access controls across the entire ML lifecycle What We Are Looking For 3-6 years of experience in software engineering, DevOps, or data platform engineering, with at least 2 years dedicated specifically to MLOps in production Strong proficiency in Python and shell scripting, alongside experience with containerization using Docker and orchestration with Kubernetes Hands-on experience with cloud infrastructure, preferably AWS or GCP, and Terraform for Infrastructure as Code (IaC) Deep familiarity with ML tracking and registry tools such as MLflow, Weights & Biases, or cloud-native model registries Solid understanding of software engineering best practices, including git workflows, unit testing, and automated integration testing Bonus: Experience with large-scale distributed training frameworks (Ray, Horovod) or deployment of Large Language Models (LLMs) using vLLM or Hugging Face TGI
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago