Site Reliability Engineer

Rackspace Technology
Posted on
Rackspace Technology logo

Experience
4 - 7 yrs
Salary (CTC)
₹11.9L - ₹17.7L
Job Location
Hyderabad, India
Vacancy
5
Designation
Site Reliability Engineer
Job Type
Not specified

Job Description

Site Reliability Engineer – L2

Shift - Should be ok for 24x7 Shift 

Location: Hyderabad - Work from Office 

Experience: 3–8 Years



About the Role
We're looking for an SRE L2 to own reliability, incident response, and automation across our production infrastructure. You'll serve as an escalation point, drive root cause analysis, and reduce toil through scripting and tooling.

Key Responsibilities

Own L2 incident response, RCA, and post-mortems for production issues.

Monitor system health and maintain SLA/SLO adherence.

Automate operational tasks to eliminate repetitive toil.

Collaborate with dev teams on deployment reliability and capacity planning.

Participate in on-call rotation and maintain runbooks.



Required Skills:

Operating Systems Hands-on with Linux (RHEL/Ubuntu) — systemd, process management, file systems, performance tuning. Working knowledge of Windows Server and event log analysis.

Cloud Practical experience on AWS / Azure / GCP — compute, storage, IAM, networking, and managed services. Familiarity with Terraform or equivalent IaC tools.

Scripting & Automation Proficiency in Python and Bash for automation, API interaction, and operational tooling. Exposure to Ansible or similar config management is a plus.

Network Troubleshooting (In-Depth) Strong command of TCP/IP internals — handshake lifecycle, connection states (TIME_WAIT, CLOSE_WAIT, SYN_FLOOD), packet flow, and socket behaviour. Hands-on with tools like tcpdump, Wireshark, netstat/ss, traceroute, mtr, and dig. Solid understanding of DNS resolution, TLS/SSL negotiation, NAT, firewalls, and routing. Able to diagnose latency, packet loss, port exhaustion, and network-level bottlenecks at the OS and infrastructure layer.

Application & HTTP Troubleshooting Deep understanding of HTTP/HTTPS methods, status codes, headers, and request lifecycle. Comfortable debugging through curl, Postman, access logs, and reverse proxy configs (Nginx / HAProxy).

Observability Experience with Prometheus, Grafana, Datadog, or ELK. Ability to build dashboards, configure meaningful alerts, and trace issues end-to-end.


Good to Have

Kubernetes / Docker experience.

Familiarity with message queues (Kafka, RabbitMQ).

Basic database troubleshooting (MySQL / PostgreSQL / Redis).

ITIL fundamentals and ITSM tools (Jira SM / ServiceNow).

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.