Lead Site Reliability Engineer
Lead Site Reliability Engineer Key Responsibilities: Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts. Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services. Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems. Automate incident detection and response using...
- AWS
- Kubernetes
- Terraform
- Linux
- Grafana
- Prometheus
- Python
- CI/CD