
Site Reliability Engineer -- KUMDC5841074
Compunnel Inc. · Canada
Remote
About the job
Job Title: Site Reliability Engineer
Experience Level: 10+ years
Location: Remote
This role is responsible for keeping production systems running, instrumenting infrastructure and application layers, building meaningful monitoring and actionable alerting, supporting incident response, and continuously improving dashboards used by engineering, operations, risk, and executive stakeholders.
Required Qualifications
• 10+ years of experience in site reliability engineering, systems engineering, software engineering, DevOps, infrastructure engineering, or production operations
• Hands-on experience with observability practices, including monitoring, alerting, logging, metrics, tracing, dashboards, and service health reporting
• Experience instrumenting applications, services, APIs, infrastructure, databases, and cloud components to enable end-to-end operational visibility
• Strong understanding of reliability engineering concepts, including SLIs, SLOs, SLAs, error budgets, incident management, capacity management, and operational readiness
• Experience designing actionable alerts that support rapid issue detection, triage, escalation, and resolution
• Experience building and maintaining operational dashboards for technical teams, support teams, and senior/executive stakeholders
• Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages for automation and operational tooling
• Experience with cloud platforms such as AWS, Azure, or GCP
• Experience with Infrastructure-as-Code tools such as Terraform or similar technologies
• Experience working with CI/CD pipelines, DevOps workflows, release processes, and production support models
• Experience troubleshooting distributed systems, REST services, event-driven architectures, messaging platforms, and service-to-service integrations
• Familiarity with relational and non-relational databases, such as PostgreSQL, MSSQL, MongoDB, or similar platforms
• Strong analytical, troubleshooting, and problem-solving skills with the ability to diagnose complex technical issues across multiple layers of the stack
• Strong written and verbal communication skills, including the ability to translate technical issues into clear business and executive-level updates
Preferred Skills
• Experience supporting cybersecurity, risk, resilience, compliance, or enterprise security platforms
• Experience with observability and monitoring tools such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, OpenTelemtry, or similar platforms
• Experience creating executive-level service health dashboards, reliability scorecards, operational risk reporting, or incident trend reporting
• Experience developing automated health checks, synthetic monitoring, service dependency maps, and operational runbooks
• Experience with incident response, major incident management, postmortems, root-cause analysis, and problem management practices
• Experience with containerized and cloud-native environments, including Kubernetes, Docker, serverless services, or managed cloud platforms
• Experience with distributed messaging or streaming platforms such as Apache Kafka
• Familiarity with cloud-native security, governance, and policy tooling such as Azure Policy, AWS SCP, GCP constraints, or related controls
• Familiarity with Cloud Security Posture Management tools such as Wiz, Prisma, Cloud Guard, or similar platforms
• Experience with cloud-based AI services such as Azure AI, AWS Bedrock, or Google Vertex AI, particularly from an operational monitoring, reliability, or governance perspective
• Experience supporting Linux and Windows environments through scripting, automation, monitoring, and operational troubleshooting
• Exposure to web technologies, APIs, front-end services, or user-facing application monitoring.
Ready to apply?Apply now