Experience
3 - 8 yrs
Salary (CTC)
₹11L - ₹14.6L
Job Location
Sahibzada Ajit Singh Nagar, India
Vacancy
1
Designation
Senior QA Engineer
Job Type
ONSITE
Job Description
Role Overview:
- The SRE/Observability Engineer owns reliability and observability for the Backend Platform: dashboards and alerting integrated with the client s OSS/OBF fault and performance management platforms, SLO practice, runbooks, incident root-cause analysis, and Kubernetes troubleshooting across the application, Kafka, and MySQL tiers.
- The role anchors the platform s path to independent ownership validating cluster and pipeline health during Shadow & Co-Own, then owning on-call and incident management (with L3/L4 Support) from Month 5. The SRE also acts as designated backup for the single DevOps/Platform Engineer.
- Own observability integration with OSS/OBF platforms: dashboards, alert rules, and fault/performance data flows.
- Build and maintain Prometheus/Grafana dashboards and alerting for all 8 services, the Kafka cluster, and the MySQL HA cluster.
- Define and operate SLI/SLO practice: availability, latency, and error-budget tracking with agreed targets.
- Author, validate, and continuously improve runbooks and MOPs; drive the documentation of tribal operational knowledge.
- Lead incident response and root-cause analysis; run blameless postmortems and track corrective actions to closure.
- Troubleshoot Kubernetes workloads, networking, and resource issues across application, messaging, and data tiers.
- Validate DB and Kafka cluster health signals during transition; tune alert thresholds to reduce noise.
- Own on-call readiness for Independent Ownership (8 5 coverage model) with structured escalation to Architects.
- Act as backup for the DevOps/Platform Engineer on Helm, CI/CD, Vault, and platform operations.
- Feed reliability improvements into the continuous-improvement backlog.
- 5+ years in SRE, production engineering, or platform operations roles.
- Strong observability tooling experience: Prometheus, Grafana, Alertmanager; log stacks (ELK/Loki).
- Proven Kubernetes troubleshooting skill: pod lifecycle, networking, storage, resource contention.
- Working SLO/error-budget practice and incident management experience (detection mitigation RCA).
- Experience writing and executing runbooks/MOPs for production systems.
- Scripting proficiency (Bash/Python) for automation and diagnostics.
- Experience monitoring Kafka and MySQL in production (lag, ISR, replication, saturation signals).
- Strong written communication for RCA reports and escalation to client stakeholders.
- CKA and/or CKS certification.
- Experience with telecom OSS/OBF-style fault and performance management platforms.
- Distributed tracing (Jaeger/OpenTelemetry) rollout experience.
- APM tools (New Relic, Dynatrace, Datadog).
- Helm/ArgoCD working knowledge sufficient to back up the DevOps function.
- Experience taking over monitoring estates from an incumbent team.
This role is based at the office Location and requires full-time, on-site presence. Candidates are expected to work from the office on all working days. Please note that this is not a work-from-home position.
Interview Process
- Assessment
- 2 Technical Rounds
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
