JJobsSonar

Site Reliability Engineer

Infojini Inc · Brazil

Remote
About the job Site Reliability Engineer 8+ years of experience LATAM What You’ll Do Incident Response & Command: · Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions · Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency · Drive root cause isolation within 30 minutes for critical incidents whenever possible · Communicate effectively across engineering, product, and leadership during high-pressure situations · Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making Proactive Reliability Engineering · Identify patterns, trends, and signals to prevent incidents before they occur · Continuously improve alert quality, reduce noise, and increase signal fidelity · Partner with engineering teams to enhance system resilience and reliability Automation & Toil Reduction · Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows · Build and improve tooling across incident response, observability, and operations · Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value Platform & Systems Support Troubleshoot across a hybrid ecosystem including: · On-prem VMs (Linux & Windows; VMware) · Cloud platforms (AWS, GCP, Azure) · Containerized environments (Kubernetes clusters) Diagnose and resolve issues across: · Networking (connectivity, latency, DB access interruptions) · Kubernetes (ingress, environment variables, cluster-level issues) · CDN and traffic management layers (Akamai, waiting rooms – plus) Required Technical Skills & Experience Core Engineering & Operations · Strong experience in incident management and triage in production environments · Proven ability to troubleshoot complex distributed systems under pressure · Solid understanding of Linux systems administration (including performance, networking, NTP, etc.) Cloud & Infrastructure · Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2) · Familiarity with GCP and/or Azure environments · Experience operating in multi-cloud and hybrid environments Containers & Orchestration · Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues) · Understanding of containerized application architectures DevOps & CI/CD Strong knowledge of DevOps practices and CI/CD pipelines Hands-on experience with: · Harness · GitHub and/or GitLab Application & Technology Stack Awareness Working knowledge of: · Java, Node.js, React-based applications Understanding of database connectivity and dependencies across: · Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required) Networking Strong foundational knowledge of: · TCP/IP, DNS, HTTP(S) · Load balancing and network troubleshooting · Diagnosing connectivity issues between services and databases Preferred Qualifications Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications Prior experience as an Incident Commander or similar leadership role during outages Familiarity with Akamai CDN and traffic management tools Experience in high-volume, high-availability production environments
Ready to apply?Apply now

Similar jobs

Browse all

Cloud Platform Engineer

ASE (Analysis Simulation Engineering) AG · Zurich, Zurich, Switzerland

On-siteEasy apply2w ago

DevOps Engineer

Sundayy · United States

RemoteEasy apply401(k), Medical2w ago