
Software Engineer, Infrastructure
jobgether · India
Accountabilities: Design, implement, and operate cloud infrastructure supporting large-scale distributed systems and data-intensive workloads. Build and manage highly available Kubernetes environments with a focus on scalability, resilience, and operational efficiency. Develop and optimize CI/CD pipelines to improve deployment reliability, testing automation, and release velocity. Enhance developer productivity by automating workflows, reducing operational friction, and improving platform tooling. Contribute to end-to-end testing frameworks and release confidence initiatives to ensure system stability. Strengthen observability capabilities through the implementation of logging, monitoring, tracing, and alerting solutions. Collaborate with engineering teams on infrastructure architecture, reliability engineering, and system scaling strategies. Support customer-facing technical integrations, troubleshoot infrastructure-related issues, and ensure reliable deployments across diverse cloud environments. Drive operational excellence by implementing infrastructure best practices, automation strategies, and reliability-focused processes. Participate in the continuous improvement of platform performance, cost optimization, and engineering efficiency. Requirements Minimum of 5 years of experience in infrastructure engineering, platform engineering, site reliability engineering, or distributed systems development. Strong programming skills in languages such as Go, Python, Java, or similar technologies. Extensive hands-on experience with Kubernetes in production environments, including cluster management and container orchestration. Strong knowledge of cloud platforms, including AWS, Google Cloud Platform (GCP), and/or Microsoft Azure. Experience with container technologies such as Docker and cloud-native application architectures. Proven expertise in designing and managing CI/CD pipelines and modern deployment practices. Strong proficiency with Infrastructure as Code (IaC) tools, particularly Terraform. Ability to troubleshoot and resolve complex cross-layer issues involving networking, storage, runtime environments, and distributed systems. Demonstrated track record of building scalable, reliable, and cost-efficient infrastructure solutions. Strong communication skills and the ability to work effectively in fast-paced, highly collaborative environments. Experience with data platforms, lakehouse architectures, or technologies such as Spark, Trino, Iceberg, Delta Lake, or Airflow is considered an advantage. Familiarity with observability and monitoring tools such as Prometheus, Grafana, ELK Stack, Datadog, or similar platforms is a plus. Strong sense of ownership, problem-solving capabilities, and a passion for infrastructure innovation and continuous improvement. Benefits Competitive compensation package, including performance-based incentives and long-term growth opportunities. Meaningful equity participation, allowing employees to contribute to and benefit from organizational success. Comprehensive healthcare coverage and employee wellbeing support. Flexible paid time off policy promoting work-life balance and personal wellbeing. Opportunity to work on cutting-edge AI infrastructure and exabyte-scale data systems. Support for continuous learning, research initiatives, conference participation, and professional development activities. Collaborative, high-impact engineering culture with significant ownership and technical autonomy. Exposure to advanced technologies and complex large-scale engineering challenges within a rapidly growing environment.
Ready to apply?Apply now