
Site Reliability Engineer
Amtex Systems Inc · United States
Remote
About the job
Site Reliability Engineer
Remote
Fulltime Opportunity
Must Have:
5+ years as SRE
Grafana – must have
PromQL, Loki, Prometheus, Tempo (These are kind of grouped, if they have 1 they should have the others) - must have
Open Telemetry, hands on experience – must have
Otel and W3c Tracing experience - must have
Python – would like to see
Docker – nice to have
Job Description:
Design and implement comprehensive SRE monitoring for web portal on GCP
Set up JVM metrics collection and performance monitoring for Java applications using GCP Monitoring
Implement logging and tracing standards across all portal components using Cloud Logging and Cloud Trace
Configure APIGEE monitoring and API performance tracking for portal services
Implement distributed tracing with W3C Trace Context headers and OpenTelemetry
Create drill-down dashboards with correlation between metrics, logs, and traces using GCP tools
Integrate GCP Monitoring, Logging, and Trace with existing Prometheus/Grafana stack
Configure GMP (Google Managed Prometheus) for enhanced metrics collection
Implement UI zero code instrumentation for frontend monitoring and traceability
Create RED (Request, Error, Duration) dashboards for Performance and Production environments
Build service health dashboards with drill-down capabilities and error message analysis
Develop and maintain SRE automation/scripts within GKE namespaces (SRE and others) for monitoring, deployment, and troubleshooting.
Ready to apply?Apply now