TymblHub

© 2026 TymblHub

Site Reliability Engineer

MSys Tech India Pvt. Ltd.
Posted on
MSys Tech India Pvt. Ltd. logo

Experience
2 - 4 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
Not specified

Job Description

Job Summary

The Senior Associate Reliability Operations role is critical in ensuring the continuous, reliable, and secure operation of our SaaS products, operating in a 24x7 support capacity. This role involves proactive monitoring, incident response, and collaboration with teams across the organization to maintain optimal service levels. The Senior Associate will participate in a rotating shift schedule to ensure high availability, rapid issue resolution, and support for key reliability initiatives. Senior Associate will serve as a key escalation point, mentor junior team members, and lead critical efforts to optimize operational workflows and systems.

Responsibilities

  • 24x7 Monitoring and Support: Oversee the health, performance, and availability of cloud-based SaaS infrastructure and applications, using monitoring tools like Prometheus and Grafana, and respond to alerts during assigned shifts.
  • Alignment and adherence to organization processes to maintain the SLA.
  • Incident Management: Act as the first responder in a 24x7 rotation, managing and mitigating service disruptions, following standard incident procedures, and escalating issues to SMEs as needed.
  • Deployments and Change Management: Manage the deployment lifecycle of the applications. Proactively engage with SMEs to resolve deployment process issues or challenges.
  • Troubleshooting and Resolution: Use diagnostic tools and scripts to resolve common issues in real-time and collaborate with cross-functional teams to analyze and address root causes.
  • Service Health and Reliability: Assist in defining and refining SLAs, SLOs, and SLIs; perform routine checks and follow established runbooks to maintain consistent service reliability.
  • Analysis and Reporting: Regularly review incident data to identify patterns, improve service resilience, and produce shift reports, summarizing system health and resolved incidents.
  • Documentation and Knowledge Base: Document incident resolutions, update runbooks, and contribute to an internal knowledge base to improve team response and efficiency.
  • Continuous Improvement Initiatives: Participate in reliability enhancement projects, including automation, configuration management, and tools improvement.
  • Collaboration: Communicate effectively with SMEs to relay critical incident information, insights, and preventive recommendations.
  • Mentorship: Work closely with team members to provide guidance during shifts and share insights on improving incident response.

Skills

  • Proficiency in monitoring and alerting tools, such as Prometheus, Grafana, Datadog, or Splunk.
  • Ability to remain composed in high-stakes situations and resolve incidents promptly.
  • Strong verbal and written communication skills to document and relay incident information effectively.
  • Linux/Unix Administration
  • Shell/Bash Scripting
  • Prometheus/Grafana Monitoring
  • Log Analysis (Splunk/Datadog/ELK)
  • Incident & Problem Management
  • Production Support (24x7 Operations)

Qualifications

  • Education: B.Sc IT, B.Sc Computers, BCA or equivalent.
  • Experience: 2-4 years of experience in reliability operations or related 24x7 support role within SaaS or cloud environments.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.