Experience
5 - 10 yrs
Salary (CTC)
₹90k - ₹1.5L
Job Location
India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
Not specified
Job Description
Role & responsibilities
- Design, implement, and maintain highly available, scalable, and reliable production systems using SRE best practices.
- Define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budget processes.
- Manage cloud-native infrastructure across AWS, Azure, or GCP with focus on reliability, performance, and automation.
- Build and maintain CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, Azure DevOps, and GitOps practices.
- Automate infrastructure provisioning, deployments, monitoring, and operational workflows using Terraform, Ansible, Python, Bash, and other automation tools.
- Manage Kubernetes-based platforms including cluster deployment, scaling, troubleshooting, upgrades, and workload optimization.
- Implement observability solutions using Prometheus, Grafana, ELK Stack, Splunk, Datadog, CloudWatch, and Azure Monitor.
- Monitor system health, analyze performance metrics, configure alerts, and proactively identify reliability risks.
- Lead incident response, production troubleshooting, root cause analysis (RCA), and post-incident improvement initiatives.
- Develop strategies for high availability, disaster recovery, fault tolerance, backup, and business continuity.
- Optimize application and infrastructure performance, scalability, and cloud resource utilization.
- Collaborate with development, DevOps, security, and infrastructure teams to improve software delivery and operational excellence.
- Implement DevSecOps practices including vulnerability management, security automation, access controls, and compliance requirements.
- Maintain infrastructure documentation, operational runbooks, architecture diagrams, and standard operating procedures.
- Participate in on-call rotations and provide production support for critical business applications.
- Mentor junior engineers and promote SRE culture, automation, and reliability engineering practices.
Preferred candidate profile
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 5 to 10 years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Production Engineering roles.
- Strong hands-on experience with AWS, Microsoft Azure, and/or Google Cloud Platform (GCP).
- Proven experience managing Kubernetes, Docker, Helm, OpenShift, and containerized production environments.
- Strong knowledge of CI/CD pipelines, GitOps, Infrastructure as Code (IaC), Terraform, Ansible, and automation frameworks.
- Excellent scripting skills in Python, Bash, Shell, or Go for automation and operational tooling.
- Strong understanding of Linux/Unix systems administration, networking, DNS, TCP/IP, HTTP, load balancing, and troubleshooting.
- Experience with monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Splunk, Datadog, New Relic, CloudWatch, and Azure Monitor.
- Strong knowledge of incident management, RCA, SLO/SLI implementation, capacity planning, performance optimization, and reliability engineering practices.
- Experience with microservices architecture, cloud-native applications, service mesh technologies, and distributed systems.
- Familiarity with DevSecOps, security best practices, IAM, vulnerability management, and compliance standards.
- Strong analytical, troubleshooting, communication, and collaboration skills.
- Experience working in Agile/Scrum environments with cross-functional engineering teams.
Preferred certifications
- Certified Kubernetes Administrator (CKA)
- AWS Certified DevOps Engineer Professional
- AWS Solutions Architect
- Microsoft Azure DevOps Engineer Expert
- Google Professional Cloud DevOps Engineer
- HashiCorp Terraform Associate
- RHCE / LFCS
- ITIL Foundation