TymblHub

© 2026 TymblHub

Site Reliability Engineer Lead

Synechron
Posted on
Synechron logo

Experience
6 - 11 yrs
Job Location
Hyderabad, India
Vacancy
4
Designation
Site Reliability Engineer Lead
Job Type
Not specified

Job Description

  • Synechron is seeking a Lead Systems Operations Engineer
  • Lead Site Reliability Engineer (App SRE) is responsible for driving reliability, automation, observability, and performance for missioncritical applications and platforms.
  • This role blends software engineering excellence with operational expertise to deliver stable, scalable, and resilient services, while reducing toil and shifting operations left across the application lifecycle.
  • The Lead SRE acts as a technical authority and mentor, partnering with application, platform, and DevOps teams to embed reliability into design, delivery, and operations.

Required Qualifications:

  • 6+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education

Job Expectations:

  • Partner with application, platform, and business stakeholders to define, implement, and govern SLIs, SLOs, and error budgets, balancing reliability with delivery velocity.
  • Lead the design and continuous improvement of observability, telemetry, monitoring, and alerting, ensuring actionable insights and reduced alert fatigue.
  • Identify, prioritize, and implement automation and selfhealing solutions to eliminate operational toil and improve service resilience.
  • Own and lead production readiness and golive activities, including NFR validation, Permit to Operate (PTO), and operational risk assessments.
  • Provide engineeringled application production support, acting as an escalation point for complex application and platform issues.
  • Lead and troubleshoot major incident response (P1/P2/P3), drive indepth root cause analysis (RCA), and ensure preventative actions are implemented to achieve longterm stability.
  • Influence and guide teams to shift reliability left by embedding SRE practices into design, CI/CD pipelines, and release processes.
  • Mentor junior engineers and contribute to SRE standards, best practices, and operating models.
  • Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals

Additional Required Qualifications:

  • 6+ years of handson experience in production application support engineering, with a strong focus on reliability, availability, and operational excellence.
  • 5+ years of experience leading and operating production systems in a Site Reliability Engineering, DevOps, or Reliability Engineering role.
  • 6+ years of experience working with enterprise schedulers and databases, such as Autosys, Oracle, and MS SQL Server.
  • 4+ years of experience supporting applications on Kubernetes / OpenShift platforms.
  • Strong understanding of observability concepts (metrics, logs, traces, APM) using tools such as AppDynamics, ThousandEyes, Prometheus, Grafana, Splunk, and Aternity.
  • Solid experience with webbased applications and application servers .
  • Proven experience providing technical leadership and handson execution in complex enterprise environments.
  • Excellent communication and documentation skills, with the ability to influence both technical and nontechnical stakeholders.