Experience
7 - 12 yrs
Job Location
Mumbai, India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
Not specified
Job Description
About the Role
We are looking for a Senior DevOps / Site Reliability Engineer (SRE) with 7+ years of experience to join our high-performing engineering team. This role is pivotal in building scalable systems, reducing operational toil, and improving the reliability of our infrastructure and applications. As a senior member, you ll lead initiatives around automation, observability, and resilience, while collaborating closely with product, development, and operations teams .
Responsibilities:
- Lead initiatives to identify and eliminate manual, repetitive tasks through automation and tooling.
- Develop self-healing infrastructure solutions and drive continuous operational efficiency.
Lead efforts to build resilient systems and proactively identify potential points of failure across the stack.
- Design and implement reliability-focused automation and tooling to ensure consistent system performance and uptime.
- Support post-release validations and operational readiness assessments to ensure smooth rollouts.
- Occasional weekend support may be required (e.g., during major releases or critical changes).
- Design, implement, and manage cloud-native infrastructure using Terraform and other IaC tools.
- Ensure infrastructure follows principles of scalability, fault tolerance, and security.
- Design and implement robust monitoring and alerting solutions using Elastic Stack, OpenTelemetry (OTEL), and similar tools.
- Define and manage SLIs/SLOs, and partner with development teams to ensure service reliability.
- Partner with engineering teams to create & improve CI/CD pipelines and deployment processes.
- Provide technical leadership and recommendations to improve system architecture, release velocity, and developer productivity.
Requirements:
- Good experience on OS - Linux, Cloud - AWS cloud
- Strong in Terraform and Ansible infrastructure
- Good experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering.
- Strong experience with AWS services and Terraform for IaC .
- Deep understanding of incident response, post-mortem analysis, and reliability engineering principles.
- Proven track record with Elastic Stack, or other observability tools.
- Proficient in scripting (Python, Bash, etc.) and working with Git-based workflows.
- Solid grasp of modern CI/CD tooling and software development lifecycle practices.
Good to Have Skills:
- Experience in Azure, Kubernetes, or container orchestration tools.
- Good to have OpenTelemetry
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
