Experience
6 - 11 yrs
Salary (CTC)
₹23.2L - ₹31.1L
Job Location
Hyderabad, India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
Not specified
Job Description
About the Role
As a Senior Site Reliability Engineer (SRE) at Kensho, you will be a hands-on technologist who combines strong infrastructure expertise with solid software engineering skills Python first. You will be responsible for ensuring the reliability, scalability, and security of both business-critical internal systems and external, customer facing services.
You will work closely with Infrastructure, Application, and Security teams to design resilient systems, automate operations, and continuously improve platform stability. This role requires deep ownership of production systems, strong troubleshooting skills across infrastructure, container orchestration systems, networking, and applications, and comfort operating in a 24/7 on call environment.
What Youll Do - Own and operate production services supporting critical financial applications with a strong focus on availability, performance, and reliability
- Design, build, and manage AWS infrastructure, including EKS-based clusters, across lower and production environments
- Provision and manage infrastructure using Terraform (Infrastructure as Code) with a strong automation first mindset
- Deploy, scale, and troubleshoot applications running on like Kubernetes, including cluster creation, upgrades, and lifecycle management
- Build and maintain automation frameworks and tooling Python based to reduce operational toil and prevent recurring incidents
- Monitor system health using metrics, logs, and alerts; continuously tune alerts, dashboards, and runbooks
- Troubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behavior
- Manage certificate lifecycle and expiration, ensuring secure and uninterrupted service operation
- Collaborate with InfoSec, Vulnerability Management, and Network Security teams (e.g., Zscaler) to maintain a strong security posture
- Collaborate with L1/L2 teams, helping them understand infrastructure and operational best practices
- Participate in on call and lead incident response, drive root cause analysis, and ensure effective post-incident remediation and learnings
- Identify architectural anti-patterns and drive improvements by reviewing new services for production readiness, resiliency, and secure design prior to release
- Establish and enforce production readiness standards, including deployment strategies, rollback plans, and observability requirements
- Optimize infrastructure cost and resource utilization without compromising reliability and performance
- 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering roles
- Strong software engineering background, with hands-on Python development used for automation, tooling, and system reliability
- Experience building or supporting scalable, distributed systems in production
- Deep experience with AWS cloud environments, including AWS, IAM, networking, and access controls
- Strong hands-on expertise with similar tools like Kubernetes (EKS preferred): cluster creation, deployments, scaling, and troubleshooting
- Solid understanding of networking fundamentals (VPCs, routing, DNS, load balancing, security groups)
- Experience with CI/CD pipelines, deployment tools, and infrastructure automation
- Working knowledge of databases and query optimization, and understanding how applications behave under load
- Familiarity with similar tools like Kafka or other messaging systems
- Comfortable conducting code reviews and participating in coding focused interviews
- Strong operational mindset with experience in incident management and oncall rotations
- Clear communicator and collaborative teammate who values documentation and knowledge sharing
- Demonstrated ownership of large-scale, production systems
- Strong examples of Python based automation or internal tooling
- Contributions to open source projects, infrastructure platforms, or reliability tooling
- Experience working closely with security and compliance teams in regulated environments
- AWS, Amazon EKS, Terraform, Jsonnet
- Similar tools like Kubernetes, Helm, CI/CD tooling
- Python (automation, tooling, reliability engineering)
- Prometheus, Grafana, logging and monitoring platforms
- PostgreSQL and other production databases
- Kafka or event driven systems
- Linux (Ubuntu or similar)
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
