TymblHub

© 2026 TymblHub

SRE / Observability Engineer

Sourcefuse Technologies
Posted on
Sourcefuse Technologies logo

Experience
3 - 8 yrs
Salary (CTC)
₹11L - ₹14.6L
Job Location
Sahibzada Ajit Singh Nagar, India
Vacancy
1
Designation
Senior QA Engineer
Job Type
ONSITE

Job Description

Role Overview:
  • The SRE/Observability Engineer owns reliability and observability for the Backend Platform: dashboards and alerting integrated with the client s OSS/OBF fault and performance management platforms, SLO practice, runbooks, incident root-cause analysis, and Kubernetes troubleshooting across the application, Kafka, and MySQL tiers.
  • The role anchors the platform s path to independent ownership validating cluster and pipeline health during Shadow & Co-Own, then owning on-call and incident management (with L3/L4 Support) from Month 5. The SRE also acts as designated backup for the single DevOps/Platform Engineer.
Key Responsibilities:
  • Own observability integration with OSS/OBF platforms: dashboards, alert rules, and fault/performance data flows.
  • Build and maintain Prometheus/Grafana dashboards and alerting for all 8 services, the Kafka cluster, and the MySQL HA cluster.
  • Define and operate SLI/SLO practice: availability, latency, and error-budget tracking with agreed targets.
  • Author, validate, and continuously improve runbooks and MOPs; drive the documentation of tribal operational knowledge.
  • Lead incident response and root-cause analysis; run blameless postmortems and track corrective actions to closure.
  • Troubleshoot Kubernetes workloads, networking, and resource issues across application, messaging, and data tiers.
  • Validate DB and Kafka cluster health signals during transition; tune alert thresholds to reduce noise.
  • Own on-call readiness for Independent Ownership (8 5 coverage model) with structured escalation to Architects.
  • Act as backup for the DevOps/Platform Engineer on Helm, CI/CD, Vault, and platform operations.
  • Feed reliability improvements into the continuous-improvement backlog.
Mandatory Qualifications:
  • 5+ years in SRE, production engineering, or platform operations roles.
  • Strong observability tooling experience: Prometheus, Grafana, Alertmanager; log stacks (ELK/Loki).
  • Proven Kubernetes troubleshooting skill: pod lifecycle, networking, storage, resource contention.
  • Working SLO/error-budget practice and incident management experience (detection mitigation RCA).
  • Experience writing and executing runbooks/MOPs for production systems.
  • Scripting proficiency (Bash/Python) for automation and diagnostics.
  • Experience monitoring Kafka and MySQL in production (lag, ISR, replication, saturation signals).
  • Strong written communication for RCA reports and escalation to client stakeholders.
Desired Qualifications:
  • CKA and/or CKS certification.
  • Experience with telecom OSS/OBF-style fault and performance management platforms.
  • Distributed tracing (Jaeger/OpenTelemetry) rollout experience.
  • APM tools (New Relic, Dynatrace, Datadog).
  • Helm/ArgoCD working knowledge sufficient to back up the DevOps function.
  • Experience taking over monitoring estates from an incumbent team.
Location for Reporting:
This role is based at the office Location and requires full-time, on-site presence. Candidates are expected to work from the office on all working days. Please note that this is not a work-from-home position.
Interview Process
  • Assessment
  • 2 Technical Rounds
Disclaimer: This job description has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.