Experience
7 - 10 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Site Reliability Engineer Lead
Job Type
Not specified
Job Description
Job Summary
Lead Site Reliability Engineer
Responsibilities
- Set the direction and long-term strategy for solving complex problems; communicate timeline, scope, risks, and the technical roadmap to leadership and stakeholders.
- Keep abreast of emerging cloud technologies, running POCs to assess suitability and value.
- Lead design and implementation of large-scale AWS architectures optimized for reliability, scalability, and cost.
- Develop and refine CI/CD pipelines and automated deployments for complex environments.
- Drive error budgets; define and track SLIs/SLOs; own organization-wide reliability targets.
- Lead deep-dive RCAs for major incidents using Datadog, Prometheus, Grafana, and ELK; run blameless postmortems.
- Establish best practices for containerization (Docker, Kubernetes) and infrastructure as code (Terraform, AWS CDK).
- Mentor, coach, and provide technical escalation for SRE/DevOps engineers; foster a learning and ownership culture.
- Direct AI/automation initiatives to improve deployment, monitoring, troubleshooting, and developer efficiency (e.g., ChatGPT, Copilot, generative AI for runbooks and incident management).
- Partner with engineering, product, and business to shape reliability strategy, incident process, architecture reviews, and roadmap.
- Ensure security, regulatory, and operational compliance across cloud architecture and automation.
- Continuously communicating timeline, scope, risks, and technical road map.
What You Need
- 7-10 years in Site Reliability, DevOps, or Cloud Engineering, including substantial leadership experience. Strong understanding of SRE principles and DevOps culture; adept at applying software engineering tools, methods, and practices. Deep AWS expertise across advanced networking, identity (IAM), and multi-region architectures; experience with large-scale, multi-tier distributed systems. Advanced development and scripting skills (Python, Go, Bash) focused on automation and troubleshooting distributed systems. Mastery of CI/CD platforms (Jenkins, GitHub Actions, Argo CD), containerization (Docker, Kubernetes), and infrastructure as code (Terraform, CloudFormation, AWS CDK). Proven track record driving SLIs/SLOs, error budgets, and reliability targets at scale, plus leading complex incident processes and blameless postmortems. Strong Unix/Linux background with deep knowledge of system internals. Experience with database technologies (SQL Server, PostgreSQL, Couchbase preferred). Effective people leader and mentor with excellent communication and stakeholder management skills; bias for execution.
