JJobsSonar

Lead Platform Infrastructure Engineer

Jobgether · United States

Remote
About the job This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Platform Infrastructure Engineer based in the United States. This role owns the architecture, deployment, and operation of a mission-critical platform designed for secure, scalable delivery within regulated environments. You will lead the infrastructure strategy that enables repeatable deployments across customer-managed and on-premises environments, ensuring reliability, security, and operational excellence. Working closely with engineering, security, data, and AI/ML teams, you will build the foundation that allows complex technology solutions to scale efficiently. The ideal candidate is a hands-on infrastructure leader who thrives on automation, security-first design, and solving complex deployment challenges. This position offers the opportunity to shape platform engineering practices from the ground up while supporting customers with demanding technical requirements. You will play a key role in making enterprise-grade deployments predictable, secure, and scalable. Accountabilities Design and own the Infrastructure-as-Code foundation that enables dedicated single-tenant and on-premises deployments with minimal customer-specific effort. Build, operate, and continuously improve containerized infrastructure using Docker, Kubernetes, and K3s for smaller customer environments. Manage core platform components, including ingress, load balancing, CI/CD pipelines, secrets management, certificates, and observability systems. Define and improve deployment workflows, environment promotion strategies, release automation, and rollback procedures to ensure reliable and repeatable operations. Partner with security teams to implement encryption, network isolation, access controls, and security hardening aligned with industry best practices. Design and optimize GPU and inference infrastructure to support growing AI workloads while maintaining cost efficiency. Establish operational practices including on-call processes, service-level objectives (SLOs), incident response procedures, and capacity planning. Create and maintain technical documentation, deployment architecture materials, and operational procedures to support customer reviews and audit readiness. Drive platform reliability improvements by monitoring performance, improving automation, and reducing operational complexity. Collaborate cross-functionally with engineering, security, data, AI/ML, and customer-facing teams to ensure successful production deployments. Requirements 8-12+ years of experience building and operating production infrastructure, including at least 2 years leading or defining direction for a platform or infrastructure function. Proven experience deploying and operating software in customer-controlled, on-premises, or highly regulated environments. Deep hands-on expertise with Docker and Kubernetes in production environments, including deployments, networking, storage, upgrades, and troubleshooting. Strong understanding of Linux systems, networking fundamentals, infrastructure automation, and configuration management practices. Extensive experience with Infrastructure-as-Code principles, including modules, state management, idempotency, and reproducible infrastructure. Demonstrated ability to build secure, reliable, and automated infrastructure solutions rather than manually managed environments. Experience with technologies such as Kubernetes, K3s, Helm, Terraform, Packer, Ansible, CI/CD platforms, Vault, SOPS, PKI/mTLS, networking tools, object storage, and observability platforms. Strong knowledge of on-premises and bare-metal infrastructure; cloud experience, particularly AWS, is beneficial. Excellent collaboration and communication skills with the ability to work effectively across technical and business teams. Bachelor’s degree in Computer Science, Information Systems, or a related technical field, or equivalent practical experience. Experience working in financial services, banking, healthcare, government, or other regulated technology environments is strongly preferred. Experience with Kubernetes network policies, namespace isolation, multi-environment rollouts, GPU scheduling, or AI workload optimization is a plus. Benefits Full-time remote opportunity based in the United States. Opportunity to lead platform infrastructure strategy within a growing technology environment. Work on complex deployment challenges involving security, automation, and large-scale infrastructure. Collaborative environment with engineering, security, AI/ML, and client-focused teams. Opportunity to influence architecture, operational practices, and long-term platform scalability. Exposure to innovative solutions supporting regulated financial institutions. Up to 20% travel may be required. Competitive compensation and benefits package. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago