Experience
4 - 5 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Site Reliability Engineer 2
Job Type
Not specified
Job Description
About the role
We are looking for a skilled Site Reliability Engineer (SRE) to join our team and help build, deploy, scale, and operate highly reliable distributed systems. You will play a key role in deploying and managing platforms across multiple regions, ensuring high availability, scalability, and performance.
Key Responsibilities :
- Deploy, manage, and scale distributed platforms across multiple geographic regions
- Design and maintain Kubernetes-based infrastructure for large-scale applications; Build and manage Helm charts for efficient and repeatable deployments
- Monitor system health using Grafana dashboards and metrics; proactively identify and resolve issues; Improve system reliability, performance, and scalability through automation and best practices
- Handle large-scale deployments and improve infrastructure for growth
- Collaborate with development teams to ensure smooth CI/CD and production readiness; Implement observability, alerting, and incident response processes; Troubleshoot production issues and perform root cause analysis
- Write and maintain run books for incident response
Required Qualifications :
- 4 5 years of experience in Site Reliability Engineering, DevOps, or similar roles Strong hands-on experience with Kubernetes in production environments
- Strong experience with infrastructure as code (Terraform, git...)
- Strong experience with AWS (eks, vpc, s3, ecr, iam role etc)
- Solid experience with Helm charts for application deployment
- Strong experience in bash scripting and tooling
- Experience with large-scale distributed systems and high-availability architectures Strong understanding of containerization, micro-services, and cloud-native ecosystems
- Experience with CI/CD pipelines and automation tools
- Good debugging and problem-solving skills in production environments
Preferred Skills:
- Proficiency in at least one of the following: Golang, Python.
- Experience building and managing Grafana dashboards and metrics
- Knowledge of monitoring and observability stacks (Prometheus, Loki, etc.)
- Experience with multi-region deployments and global infrastructure