Experience
4 - 6 yrs
Job Location
Hyderabad, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
Not specified
Job Description
Key Responsibilities
Reliability Engineering & Operations
- Design, implement, and continuously improve the reliability, availability, and performance of critical systems (batch, APIs, integrations, and customer-facing platforms)
- Define and operationalize SLIs, SLOs, and error budgets for critical services in partnership with engineering and product teams
- Participate in on-call rotations, incident response, and major incident management
- Lead and contribute to blameless post-incident reviews, driving root cause analysis and measurable reliability improvements
- Proactively identify reliability risks and lead remediation efforts before they impact clients
Observability & Monitoring
- Build and maintain end-to-end observability across applications, infrastructure, and integrations (metrics, logs, traces, alerts)
- Implement actionable monitoring and alerting to reduce noise and improve signal quality
- Partner with application teams to instrument services using best-in-class observability practices
- Ensure visibility into system health, capacity, performance, and failure modes across environments
- Hands-on experience with Grafana, Azure Log Analytics, Prometheus, Datadog, Splunk, or equivalents.
Automation & Toil Reduction
- Identify repetitive operational tasks and automate them through code
- Improve deployment reliability through automation, self-service tooling, and safe rollout patterns
- Reduce manual intervention in batch processing, integrations, and operational workflows
- Apply Infrastructure-as-Code and configuration automation to improve consistency and repeatability
Cloud, Platform & Infrastructure Reliability
- Support reliability of Azure-based infrastructure, containerized workloads, and hybrid environments
- Partner with platform, DevOps, and infrastructure teams to improve resilience, scalability, and recovery
- Contribute to capacity planning, performance tuning, and cost-aware reliability decisions
- Ensure systems meet RTO/RPO, backup, and disaster recovery expectations
Secure & Compliant Operations
- Embed security, compliance, and risk controls into operational practices
- Work closely with Security and Compliance teams to meet financial services regulatory requirements
- Ensure production systems follow least privilege, secure configuration, and auditability standards
- Support vulnerability remediation and secure operational processes
- Deep understanding of environment reliability and stability engineering needs in High Availability, Fault Tolerance, Capacity Planning, Performance Engineering, Disaster Recovery, Multi-Region Architecture, Scalability, and Chaos Engineering
Collaboration & Enablement
- Partner with application engineering teams to improve production readiness and operational maturity
- Influence system design by advocating for reliability-first architectural decisions
- Provide guidance on operational best practices, deployment safety, and observability standards
- Document operational patterns, runbooks, and reliability guidelines in Confluence
- Act as a reliability advocate across engineering teams
- Strong software engineering skills in .NET / C#, Python, Java, NodeJS, or similar
- Experience operating distributed systems in production
- Deep understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, incident management
- Experience with Azure (or AWS/GCP), including compute, networking, and managed services
- Knowledge of containerization and orchestration (Docker, Kubernetes preferred)
- Experience with monitoring, logging, tracing, and alerting tools
- Familiarity with CI/CD pipelines, automation, and Infrastructure-as-Code
- Deep understanding of system administration, networking (TCP/IP, DNS, HTTPS, SSL / TLS), Storage, Load Balancing, and File Systems
- Understanding of security best practices in regulated enterprise environments
- Experience with JIRA, PagerDuty, or equivalent
- Experience supporting financial services or highly regulated systems (preferred)
- Bachelors degree in computer science, Software Engineering, or related technical field
- 4-6 years of software engineering experience of experience in Site Reliability Engineering, DevOps, Platform Engineering, or production operations
- Proven experience in troubleshooting and improving production system reliability
- Experience supporting 24/7 systems, batch processing, and mission-critical workloads
- Strong collaboration skills across engineering, security, and infrastructure teams
- Experience working in Agile/Scrum environments
- Experience building APIs, services, and/or platform components
- Understanding of enterprise integration patterns, service-oriented architecture, and large-scale system design
- Experience with DevOps practices, cross-functional collaboration, and agile/scrum development methodologies