TymblHub

© 2026 TymblHub

Chief Manager - Site Reliability Engineer

Bajaj Financial Securities
Posted on
Bajaj Financial Securities logo

Experience
5 - 10 yrs
Job Location
Pune, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE

Job Description

Job Description for Chief Manager - Site Reliability Engineer at Bajaj Broking:

As the Chief Manager - Site Reliability Engineer at Bajaj Broking, you will play a crucial role in ensuring the availability, performance, and reliability of our services. You will lead a team of engineers focused on building and maintaining the infrastructure that supports our applications and systems. Your responsibilities will include:

- Designing, implementing, and managing scalable and reliable systems that support our trading platform and associated services.
- Collaborating with development teams to ensure that application designs support operational requirements, including scalability and performance monitoring.
- Developing and implementing automation tools and frameworks for deploying and managing applications in production environments.
- Establishing and maintaining service level objectives (SLOs) and service level agreements (SLAs) to ensure that performance metrics are aligned with business goals.
- Monitoring systems to proactively identify performance bottlenecks and resolve issues before they impact users.
- Leading incident response efforts, conducting root cause analyses, and implementing preventative measures to enhance overall system reliability.
- Researching and implementing new technologies that enhance our reliability and operational efficiency.
- Mentoring and guiding team members to foster a culture of continuous improvement and learning within the Site Reliability Engineering team.

Skills Required:

- Strong experience in system administration, networking, and cloud infrastructure (AWS, Azure, or GCP).
- Proficiency in programming and scripting languages such as Python, Go, Bash, or similar languages.
- Deep understanding of container orchestration technologies (Docker, Kubernetes).
- Experience with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, or similar.
- Familiarity with configuration management tools (Ansible, Puppet, Chef).
- Knowledge of database management (SQL and NoSQL databases).
- Strong problem-solving skills and the ability to troubleshoot complex systems and applications.
- Excellent communication and leadership skills to foster collaboration among cross-functional teams.

Tools Required:

- Cloud Platforms: AWS, Azure, Google Cloud Platform
- Containerization: Docker, Kubernetes
- Monitoring and Logging: Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana)
- Configuration Management: Ansible, Puppet, Chef
- Programming/Scripting: Python, Go, Bash
- Databases: MySQL, PostgreSQL, MongoDB, Redis

This position offers an exciting opportunity to be at the forefront of technology within the financial services sector, contributing to the reliability and performance of our key trading systems.
About the Role:
- The Chief Manager - Site Reliability Engineer at Bajaj Broking will focus on enhancing the reliability and performance of our financial services platforms.
- You will lead initiatives to improve site reliability, ensuring our systems are efficient, resilient, and scalable.
- The role involves strategic planning, implementation of best practices, and oversight of incident management processes.

About the Team:
- You will be part of a dynamic team of SREs and engineers dedicated to delivering high-quality services and solutions.
- The team collaborates closely with development, operations, and support teams to foster a culture of reliability and performance.
- Emphasizing continuous learning and improvement, the team regularly engages in knowledge-sharing sessions and technical discussions.

You are Responsible for:
- Leading a team of engineers to design, implement, and maintain reliable and scalable systems.
- Managing incident response, root cause analysis, and post-mortem reviews to minimize downtime.
- Monitoring system performance, identifying bottlenecks, and proposing enhancements to improve efficiency.

To succeed in this role you should have the following:
- Strong experience in site reliability engineering, DevOps practices, and cloud infrastructure.
- Proficiency in scripting and programming languages to automate processes and improve system efficiency.
- Excellent problem-solving skills, with the ability to work under pressure and handle critical incidents effectively.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.