JJobsSonar

Senior Platform Engineer

Quantiphi · United States

Remote
About the job Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services and Solutions that help organizations solve what truly matters. We partner with enterprises to reimagine their businesses through intelligent, scalable, and transformative AI driving measurable outcomes at the very core of their operations. Since our founding in 2013, Quantiphi has tackled some of the world’s most complex business challenges by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our work is rooted in delivering accelerated, quantifiable business value, not just technology for technology’s sake. Headquartered in Boston, Quantiphi is a global organization with 4,000+ professionals serving clients across key industry verticals, including BFSI, Healthcare & Life Sciences, CPG, MFG, TME etc. As an Elite and Premier partner to leading cloud and AI platforms such as NVIDIA, Google Cloud, AWS, and Snowflake, we build and deliver enterprise-grade AI services and solutions that create real-world impact. Our industry recognition includes: 21x Google Cloud Partner of the Year awards in the last 8 years. 3x AWS AI/ML award wins. 3x NVIDIA Partner of the Year titles. 2x Snowflake Partner of the Year awards. Top analyst recognitions from Gartner, ISG, and Everest Group. Consecutive certifications as a Great Place to Work. Be part of a trailblazing team that’s shaping the future of AI, ML, and cloud innovation. Your next big opportunity starts here! For more details, visit: Website or LinkedIn Page. Role: Senior Platform Engineer Experience Level: 7 + years Work Location: Preferred New Jersey or anywhere in USA Role Overview We are looking for a Senior Platform Engineer who has built and operated production-grade AI/ML platforms on AWS — with deep, hands-on expertise in Amazon SageMaker (Studio, Training, Endpoints, Pipelines, Model Registry) and end-to-end AI/ML governance. You will own the platform backbone that GenAI, Agentic AI, and classical ML workloads run on for some of Quantiphi's most strategic BFSI accounts — with a relentless focus on data lineage, model governance, security, cost, and reliability across the full AI/ML lifecycle. Responsibilities Lead the architecture, build-out, and day-to-day operations of the enterprise AI/ML platform on AWS — with SageMaker Studio as the primary developer and data-science surface for Insurance and Financial Services clients. Design and enforce AI/ML governance in SageMaker Studio — domains, user profiles, spaces, IAM roles, lifecycle configs, VPC/PrivateLink isolation, KMS encryption, and least-privilege access aligned to BFSI regulatory requirements. Own end-to-end data and model lineage across the AI/ML lifecycle — from raw ingest through feature engineering, training, evaluation, deployment, and inference — using SageMaker ML Lineage Tracking, SageMaker Model Cards, Amazon DataZone / SageMaker Catalog, AWS Glue Data Catalog, and OpenLineage. Design, deploy, and operate SageMaker endpoints (Real-time, Serverless, Asynchronous, and Multi-Model) — including autoscaling, shadow/canary deployments, A/B routing, and cost/latency optimization for regulated workloads. Build and maintain reusable MLOps templates using SageMaker Pipelines, SageMaker Projects, Model Registry, Model Monitor, Clarify, and CI/CD (CodePipeline / GitHub Actions) — so data scientists ship models to production in days, not months. Implement platform-level guardrails for model risk management (MRM) — bias/explainability with Clarify & SHAP, drift monitoring, approval workflows, model cards, and audit trails for SOX / NAIC / GLBA / NYDFS / HIPAA obligations. Stand up and govern feature stores, vector stores, and knowledge bases (SageMaker Feature Store, OpenSearch, Bedrock Knowledge Bases) with clear ownership, lineage, and access controls. Partner with Security, Cloud, and Data Engineering teams to embed AI/ML workloads into the client's landing zone — VPC design, PrivateLink, service control policies, KMS, Secrets Manager, and network segmentation. Instrument the platform for observability — CloudWatch, CloudTrail, OpenTelemetry, model metrics, cost & usage reporting — and drive FinOps for SageMaker training and inference spend. Provide technical leadership and mentorship to ML engineers, data scientists, and junior platform engineers; drive best practices in reproducibility, lineage, and responsible AI. Serve as the trusted platform advisor to CIO / CDO / Head-of-MLOps stakeholders — translating governance, lineage, and compliance requirements into concrete AWS platform patterns. Skills Required 7+ years of hands-on platform / MLOps / cloud engineering experience, with recent, demonstrable delivery of production AI/ML platforms on AWS for Insurance or Financial Services clients. MANDATORY: Active AWS certification — AWS Certified Machine Learning – Specialty, AWS Certified Machine Learning Engineer – Associate, or AWS Certified Solutions Architect – Professional strongly preferred. Deep, hands-on expertise with Amazon SageMaker Studio — domains, user profiles, spaces, JupyterLab/Code Editor apps, lifecycle configurations, custom images, and VPC-only mode. Proven experience implementing AI/ML governance in SageMaker Studio — IAM role scoping, SCPs, KMS, network isolation, tagging, cost allocation, and audit-grade access controls for regulated BFSI workloads. Deep understanding of data and model lineage across the AI/ML lifecycle — SageMaker ML Lineage Tracking, SageMaker Model Cards, Amazon DataZone / SageMaker Catalog, AWS Glue Data Catalog, Lake Formation, and OpenLineage / Marquez. Expert-level, production experience with SageMaker endpoints — Real-time, Serverless, Asynchronous, and Multi-Model — including autoscaling, shadow deployments, blue/green & canary rollouts, and inference cost/latency tuning. Strong command of SageMaker Pipelines, Projects, Model Registry, Model Monitor, Clarify, Feature Store, and JumpStart. Heavy AWS experience across compute, data, and security services — EC2, EKS, Lambda, Step Functions, S3, Glue, Athena, Redshift, Lake Formation, IAM, KMS, VPC, PrivateLink, Secrets Manager, CloudTrail, CloudWatch. Expert-level Python programming — clean, tested, production-grade code; strong software engineering fundamentals (Git, code review, TDD). Strong Infrastructure-as-Code skills — Terraform and/or AWS CDK; hands-on with CI/CD (CodePipeline, CodeBuild, GitHub Actions, Jenkins). Solid grounding in containerization (Docker) and orchestration (Kubernetes / EKS) for AI/ML workloads. Working knowledge of GenAI / LLM operations — Amazon Bedrock, Bedrock Agents & Knowledge Bases, guardrails, and evaluation — and how to govern them alongside classical ML. Excellent communication skills — able to explain governance, lineage, and platform trade-offs clearly to data scientists, security officers, model-risk teams, and CIO/CDO-level executives. Must be based in the US and able to work Eastern Time (EST) hours; New Jersey / NYC-metro location strongly preferred. Nice to Have Experience with insurance / financial data domains — ACORD, ISO, claims/underwriting workflows, or financial document taxonomies (10-K, 10-Q, loan tapes, credit memos). Hands-on with Amazon DataZone, SageMaker Unified Studio, or emerging AWS AI governance services. Experience implementing Model Risk Management (SR 11-7 / NAIC MRM) tooling and workflows on AWS. Experience with policy-as-code (OPA / Cedar) and data-access governance (Lake Formation LF-Tags, row/column-level security). Prior consulting or client-facing delivery experience; comfort presenting to CIO/CDO/CISO-level executives. Contributions to open source (SageMaker, OpenLineage, Kubeflow, MLflow), patents, or publications on MLOps / AI governance. Why Quantiphi Own the platform backbone for category-defining GenAI, Agentic AI, and classical ML programs at Fortune 500 Insurance and Financial Services clients. Direct access to AWS product teams, early previews of SageMaker and Bedrock capabilities, and Quantiphi's proprietary baioniq platform. A culture of technical depth, ownership, and rapid growth — with a clear architect track from Senior Platform Engineer to Principal / Platform Architect.
Up to $160K/yrApply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago