Experience
8 - 13 yrs
Job Location
Pune, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE
Job Description
Manager, Site Reliability Engineering
Location: Pune - Hybrid
Position Summary:
SRE Managers operate with significant autonomy and are accountable for the operational success of an entire service area, product platform, or collection of services. They balance people leadership, technical leadership, and operational ownership, regularly influencing priorities beyond their immediate team and partnering with engineering and product leaders to improve reliability, operational efficiency, and customer outcomes.
They are trusted to make decisions that affect multiple teams and are expected to build systems, processes, and organizations that scale.
Responsibilities:
The Manager, SRE leads one or more teams of Site Reliability Engineers and is accountable for reliability, availability, operational excellence, automation strategy, and engineering effectiveness within their assigned service area. Responsibilities span four primary domains:
Reliability Leadership
- Own availability, performance, scalability, and operational health of assigned services.
- Drive adoption and enforcement of SRE practices including SLOs, SLIs, error budgets, incident management, and operational reviews.
- Own reliability roadmaps and ensure investment is balanced between feature delivery, technical debt reduction, and operational sustainability.
- Establish measurable operational targets and continuously improve system resilience.
- Strategize and communicate with management of Support teams on releases, automation changes, SOPs, and ticket management.
Engineering People Leadership
- Build a high-performing team through hiring, coaching, mentoring, and career development.
- Set clear expectations and performance standards aligned to organizational objectives.
- Develop future technical and people leaders within the team.
- Ensure engineering excellence through design reviews, operational reviews, and technical governance.
- Collaborate with management of Development, QA, Product, and Release teams to ensure SRE principles, defensive coding practices, and release readiness posture are shifted left into planning, design, development, and delivery processes.
Operational Excellence
- Lead response and recovery efforts for critical incidents.
- Drive systemic problem resolution through blameless postmortems and automation.
- Eliminate toil through software engineering and platform capabilities.
- Ensure operational readiness for new products, services, and major releases.
Strategic Execution Organizational Leadership
- Translate organizational strategy into actionable roadmaps and quarterly objectives.
- Work across development, product management, cloud operations, security, and support teams to align priorities.
- Guide architectural decisions that improve reliability, scalability, security, and cost efficiency.
- Lead multiple concurrent initiatives spanning products, platforms, or service areas.
- Establish operating mechanisms, standards, and best practices across teams.
- Represent SRE during executive reviews, customer escalations, and strategic planning discussions.
- Participate in budget planning, workforce planning, and organizational design discussions.
Qualifications:
- Bachelor s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
- Typically 8+ years of experience across software engineering, cloud infrastructure, platform engineering, systems administration, or reliability engineering.
- Minimum 4 years of experience with:
- Public cloud platforms (Azure, AWS, or GCP)
- Infrastructure as Code (Terraform, Bicep, or Ansible)
- Kubernetes and containerized workloads
- Observability tooling and telemetry systems
- CI/CD and software delivery practices
- Minimum 3 years of experience leading engineers through formal people-leadership, technical leadership, or team management responsibilities.
- Experience operating and supporting business-critical production services.
- Experience leading major incidents, postmortems, and reliability improvement programs
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
