Job Description
Location: Hyderabad, India
Incident Manager About the RoleThe Incident Manager is responsible for leading the end-to-end lifecycle of major and high-impact incidents to restore service as quickly as possible while minimizing business impact. This role drives structured incident response, coordinates cross-functional teams, ensures clear executive communication, and participates in post-incident reviews to prevent recurrence.
This role partners closely with SRE, Infrastructure, Product, Engineering, Vendor Partner, and Business teams to minimize customer impact, reduce Mean Time to Restore (MTTR), and strengthen system resilience through data-driven continuous improvement.
The Incident Manager is a high visibility role with executive exposure. This role participates in an on-call rotation and must be available to manage incidents outside normal business hours as required.
The Incident Manager acts as the operational conductor during incidents and as a reliability advocate outside of them.
What You ll Do
Incident Management
- Lead and coordinate Incident bridges, Major Incident bridges, and command centers.
- Rapidly assess business impact and determine severity.
- Mobilize appropriate product, engineering, vendor partner, and business teams.
- Drive service restoration with grace and the appropriate urgency.
- Manage escalations when required.
- Ensure accurate, timely communication with executives and impacted stakeholders.
Incident Process Governance
- Continuously improve the Incident Management process.
- Ensure adherence to SLAs, SLOs, and defined severity criteria.
- Maintain runbooks, response procedures, and escalation matrices.
- Conduct quality reviews of incident records and timelines.
- Lead and participate in BCP, DR, and Incident Table-Top exercises.
Communication & Stakeholder Management
- Provide clear, concise executive-level summaries during active incidents.
- Coordinate internal and external customer communications.
- Serve as the single point of accountability during major incidents.
- Facilitate collaboration across product, engineering, infrastructure, security, crisis management, business, and vendor partner teams.
Post-Incident Management
- Participate in blameless post-incident reviews (PIRs).
- Ensure root cause analysis (RCA) is completed and documented.
- Track corrective and preventive actions to closure.
- Identify systemic themes and trends for leadership reporting.
Metrics & Continuous Improvement
- Track and report on KPIs such as:
- MTTR (Mean Time to Restore)
- Incident volume by severity
- Recurring incident trends
- SLA adherence
- SLO Adherence
- Identify opportunities for automation and monitoring improvements.
- Partner with SRE and Product Engineering teams to reduce incident frequency and duration.
What You Bring
Required Experience & Skills
- 5+ years of experience in IT Operations, SRE, or Incident Management.
- Demonstrated experience leading high-severity production incidents.
- Strong understanding of ITIL Incident and Problem Management processes.
- Exceptional facilitation and communication skills under pressure.
- Experience working with cross-functional teams.
- Ability to influence without direct authority.
- Strong analytical and organizational skills.
Preferred Qualifications
- ITIL v3 or v4 certification.
- Experience with incident management platforms (e.g., ServiceNow, Jira Service Management).
- Familiarity with on-call management tools (e.g., PagerDuty).
- Experience in regulated or high-availability enterprise environments.
- Understanding of cloud platforms (e.g., Amazon Web Services, Microsoft Azure, Google Cloud Platform).
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.