TymblHub

© 2026 TymblHub

Lead Data Centre

ACT Fibernet
Posted on
ACT Fibernet logo

Experience
4 - 6 yrs
Salary (CTC)
₹6L - ₹10L
Job Location
Hyderabad, India
Vacancy
1
Designation
Lead Data Engineer
Job Type
ONSITE

Job Description

Job Title:


Lead Site Reliability Engineering (SRE)


Role Summary:


Leading the site reliability and infrastructure operations function, responsible for ensuring high availability, scalability, and performance of hosted websites and cloud infrastructure. Owns end-to-end management of Kubernetes clusters, cloud environments, monitoring systems, incident management frameworks, and infrastructure automation. Acts as a subject matter expert in ITIL practices, configuration management, and automation feasibility to drive operational excellence.


Key Responsibilities:


 Manage and maintain Kubernetes clusters hosting production websites, ensuring reliability, scalability, and performance.

 Administer cloud accounts across AWS, Azure, and GCP, including user access management and cost optimization initiatives.

 Own the monitoring and observability stack end-to-end using Prometheus and Grafana — building dashboards, configuring alerts, and driving proactive issue detection.

 Incident Management & Response: Lead major incident management (MIM) processes, root cause analysis (RCA), and post-mortems (blameless) using enterprise ITSM tools like ServiceNow or BMC Remedy to minimize MTTR and maintain strict SLA/SLO standards.

 CMDB & Asset Management: Maintain strong working knowledge of Configuration Management Databases (CMDB) within ServiceNow/BMC Remedy to ensure accurate mapping of IT infrastructure, application dependencies, and service relationships.

 Automation Feasibility & IaC: Design and implement end-to-end infrastructure automation using Terraform (Infrastructure as Code) and Ansible (configuration management). Continuously assess legacy and operational processes to evaluate automation feasibility and eliminate toil.

 Build and maintain CI/CD pipelines to enable reliable, automated deployments.

 Manage containerized applications and microservices architecture using Docker and Kubernetes.

 Drive best practices for infrastructure reliability, security, and cost efficiency across multi-cloud environments.

 Collaborate with engineering teams to improve system reliability, incident response, and operational efficiency.


Technical Skills:


Unix/Linux, Incident Management, SDLC, Kubernetes, Docker, Microservices, Terraform, Ansible, Infrastructure as Code (IaC), Automation Feasibility Assessment, Incident Management, Service Desk / ITSM Tools (ServiceNow, BMC Remedy), CMDB Understanding & Dependency Mapping, CI/CD Pipelines, AWS, Azure, GCP, Prometheus, Grafana, Prometheus, Cloud Cost Optimization.


be ready for Rotational Shifts and alternative Saturday working (2nd and 4th Sat off) rest all Saturdays working.





No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.