Lead SRE

Sigma Allied Services private Limited
Posted on

Experience
4 - 8 yrs
Job Location
Hyderabad, India
Vacancy
1
Designation
Site Reliability Engineer Lead
Job Type
Not specified

Job Description

We are expanding our Site Reliability Engineering (SRE) organization to enable a 24x7 follow-the-sun operating model. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.

This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters our SREs act decisively to prevent service degradation and protect the customer experience.

What You ll Do

  • Incident Response & Command
  • Lead team of 5/6 SREs and works closely with client Leads and Managers for daily activities
  • Drives Automation to improve Performance of the systems
  • Prepare Weekly/ Monthly KPI reports
  • Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions
  • Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency
  • Drive root cause isolation within 30 minutes for critical incidents whenever possible
  • Communicate effectively across engineering, product, and leadership during high-pressure situations
  • Maintain a strong presence on incident bridges this role requires confidence, ownership, and clear decision-making

Proactive Reliability Engineering

Identify patterns, trends, and signals to prevent incidents before they occur

Continuously improve alert quality, reduce noise, and increase signal fidelity

Partner with engineering teams to enhance system resilience and reliability

Automation & Toil Reduction

Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows

Build and improve tooling across incident response, observability, and operations

Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value

Platform & Systems Support

Troubleshoot across a hybrid ecosystem including:

On-prem VMs (Linux & Windows; VMware)

Cloud platforms (AWS, GCP, Azure)

Containerized environments (Kubernetes clusters)

Diagnose and resolve issues across:

Networking (connectivity, latency, DB access interruptions)

Kubernetes (ingress, environment variables, cluster-level issues)

CDN and traffic management layers (Akamai, waiting rooms plus)

Required Technical Skills & Experience

Core Engineering & Operations

Strong experience in incident management and triage in production environments

Proven ability to troubleshoot complex distributed systems under pressure

Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)

Cloud & Infrastructure

Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)

Familiarity with GCP and/or Azure environments

Experience operating in multi-cloud and hybrid environments

Containers & Orchestration

Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues)

Understanding of containerized application architectures

DevOps & CI/CD

Strong knowledge of DevOps practices and CI/CD pipelines

Hands-on experience with:

Harness

GitHub and/or GitLab

Application & Technology Stack Awareness

Working knowledge of:

Java, Node.js, React-based applications

Understanding of database connectivity and dependencies across:

Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)

Networking

Strong foundational knowledge of:

TCP/IP, DNS, HTTP(S)

Load balancing and network troubleshooting

Diagnosing connectivity issues between services and databases

Preferred Qualifications

  • Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications
  • Prior experience as an Incident Commander or similar leadership role during outages
  • Familiarity with Akamai CDN and traffic management tools
  • Experience in high-volume, high-availability production environments

Key Traits for Success

  • Bias for action you move fast and decisively when systems are at risk
  • Strong communicator especially under pressure on incident bridges
  • Systems thinker able to connect the dots across complex architectures
  • Automation mindset constantly looking to reduce toil and improve efficiency
  • Continuous learner stays current with tools, AI capabilities, and emerging technologies

Why This Role Matters

This team forms the backbone of our global reliability strategy, ensuring continuous coverage and rapid response across all hours. You will directly impact uptime, customer experience, and operational excellence playing a critical role in preventing and resolving issues before they become major incidents.

Working hours -

Shift 1 - 3 AM to 12:30 PM IST ( 5:30 PM to 3 AM EST)

Shift 2 - 11 AM to 8:30 PM IST ( 01:30 AM to 11 AM EST)

Disclaimer: This job description has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.