TymblHub

© 2026 TymblHub

Lead SRE

Cvent India Pvt. Ltd.
Posted on
Cvent India Pvt. Ltd. logo

Experience
7 - 10 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Site Reliability Engineer Lead
Job Type
Not specified

Job Description

Job Summary

Lead Site Reliability Engineer

Responsibilities

  • Set the direction and long-term strategy for solving complex problems; communicate timeline, scope, risks, and the technical roadmap to leadership and stakeholders.
  • Keep abreast of emerging cloud technologies, running POCs to assess suitability and value.
  • Lead design and implementation of large-scale AWS architectures optimized for reliability, scalability, and cost.
  • Develop and refine CI/CD pipelines and automated deployments for complex environments.
  • Drive error budgets; define and track SLIs/SLOs; own organization-wide reliability targets.
  • Lead deep-dive RCAs for major incidents using Datadog, Prometheus, Grafana, and ELK; run blameless postmortems.
  • Establish best practices for containerization (Docker, Kubernetes) and infrastructure as code (Terraform, AWS CDK).
  • Mentor, coach, and provide technical escalation for SRE/DevOps engineers; foster a learning and ownership culture.
  • Direct AI/automation initiatives to improve deployment, monitoring, troubleshooting, and developer efficiency (e.g., ChatGPT, Copilot, generative AI for runbooks and incident management).
  • Partner with engineering, product, and business to shape reliability strategy, incident process, architecture reviews, and roadmap.
  • Ensure security, regulatory, and operational compliance across cloud architecture and automation.
  • Continuously communicating timeline, scope, risks, and technical road map.

What You Need

  • 7-10 years in Site Reliability, DevOps, or Cloud Engineering, including substantial leadership experience. Strong understanding of SRE principles and DevOps culture; adept at applying software engineering tools, methods, and practices. Deep AWS expertise across advanced networking, identity (IAM), and multi-region architectures; experience with large-scale, multi-tier distributed systems. Advanced development and scripting skills (Python, Go, Bash) focused on automation and troubleshooting distributed systems. Mastery of CI/CD platforms (Jenkins, GitHub Actions, Argo CD), containerization (Docker, Kubernetes), and infrastructure as code (Terraform, CloudFormation, AWS CDK). Proven track record driving SLIs/SLOs, error budgets, and reliability targets at scale, plus leading complex incident processes and blameless postmortems. Strong Unix/Linux background with deep knowledge of system internals. Experience with database technologies (SQL Server, PostgreSQL, Couchbase preferred). Effective people leader and mentor with excellent communication and stakeholder management skills; bias for execution.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.