TymblHub

© 2026 TymblHub

Site Reliability Engineer

SPARIX GLOBAL PRIVATE LIMITED
Posted on

Experience
2 - 6 yrs
Job Location
Noida, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE

Job Description

Job Summary (Site Reliability Engineer Hybrid, Bangalore, 7:30 AM 5 PM)

- Ensure service reliability, scalability, and efficiency by applying software engineering to operations.
- Define and monitor Service Level Indicators/Objectives (SLIs/SLOs) and take action when SLOs are at risk.
- Lead incident management activities, including on-call rotations, incident response, root cause analysis, postmortems, and implementing corrective actions.
- Automate deployment, remediation, scaling, and routine runbook tasks to minimize manual efforts.
- Design, implement, and maintain observability tools such as metrics, logging, and tracing for rapid diagnosis and capacity management.
- Conduct performance testing, capacity planning, and tuning to meet system targets.
- Enhance service resilience through canary releases, chaos testing, circuit breakers, and progressive rollouts.
- Collaborate with development teams to productionize new features, strengthen service reliability, and reduce operational risks.
- Create and maintain documentation (runbooks), conduct reliability reviews, and coach teams on SRE best practices.
- Utilize strong programming/scripting skills (Python, Go, etc.), Linux, networking, and debugging expertise.
- Work with cloud platforms (AWS, Azure, GCP), containers, orchestration (Kubernetes), and Infrastructure as Code (Terraform/ARM).
- Demonstrate experience in incident leadership, postmortem practices, and operational automation (CI/CD pipelines, deployment, infrastructure automation).
- Exhibit excellent communication, calmness under pressure, collaboration, and ability to drive cross-team improvements.
- Preferred: Bachelor s degree in Computer Science/Engineering (or equivalent), prior SRE or systems engineering experience, certifications in cloud/Kubernetes/reliability engineering.
- Success measured by SLO attainment, reduced detection/recovery times, automation-driven toil reduction, closure of postmortem items, and adherence to error budget governance. Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.