Senior Site Reliability Engineer 1

REED ELSEVIER INDIA (a part of RELX India Pvt Ltd)
Posted on
REED ELSEVIER INDIA (a part of RELX India Pvt Ltd) logo

Experience
7 - 12 yrs
Job Location
Mumbai, India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
Not specified

Job Description

About the Role
We are looking for a Senior DevOps / Site Reliability Engineer (SRE) with 7+ years of experience to join our high-performing engineering team. This role is pivotal in building scalable systems, reducing operational toil, and improving the reliability of our infrastructure and applications. As a senior member, you ll lead initiatives around automation, observability, and resilience, while collaborating closely with product, development, and operations teams .
Responsibilities:
  • Lead initiatives to identify and eliminate manual, repetitive tasks through automation and tooling.
  • Develop self-healing infrastructure solutions and drive continuous operational efficiency.
    Lead efforts to build resilient systems and proactively identify potential points of failure across the stack.
  • Design and implement reliability-focused automation and tooling to ensure consistent system performance and uptime.
  • Support post-release validations and operational readiness assessments to ensure smooth rollouts.
  • Occasional weekend support may be required (e.g., during major releases or critical changes).
  • Design, implement, and manage cloud-native infrastructure using Terraform and other IaC tools.
  • Ensure infrastructure follows principles of scalability, fault tolerance, and security.
  • Design and implement robust monitoring and alerting solutions using Elastic Stack, OpenTelemetry (OTEL), and similar tools.
  • Define and manage SLIs/SLOs, and partner with development teams to ensure service reliability.
  • Partner with engineering teams to create & improve CI/CD pipelines and deployment processes.
  • Provide technical leadership and recommendations to improve system architecture, release velocity, and developer productivity.
Requirements:
  • Good experience on OS - Linux, Cloud - AWS cloud
  • Strong in Terraform and Ansible infrastructure
  • Good experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering.
  • Strong experience with AWS services and Terraform for IaC .
  • Deep understanding of incident response, post-mortem analysis, and reliability engineering principles.
  • Proven track record with Elastic Stack, or other observability tools.
  • Proficient in scripting (Python, Bash, etc.) and working with Git-based workflows.
  • Solid grasp of modern CI/CD tooling and software development lifecycle practices.
Good to Have Skills:
  • Experience in Azure, Kubernetes, or container orchestration tools.
  • Good to have OpenTelemetry

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.