Experience
3 - 7 yrs
Job Location
Kolkata, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE
Job Description
InOrg Global is looking for Site Reliability Engineer (SRE) to join our dynamic team and embark on a rewarding career journey. Service Reliability : Ensure that web services and applications are highly available, performant, and reliable. This involves minimizing downtime and service disruptions. Automation : Develop automation tools and scripts to streamline operations, deployment, and monitoring processes, reducing manual intervention and the potential for human error. Incident Management : Respond to incidents and outages promptly, working to resolve issues and prevent them from recurring. Implement incident response best practices. Service-Level Objectives (SLOs) and Service-Level Indicators (SLIs) : Define and track SLOs and SLIs to measure and improve service reliability. Use data-driven insights to set and meet reliability targets. Monitoring and Alerting : Set up monitoring systems to continuously assess the health and performance of systems. Configure alerts to notify the team when issues arise. Capacity Planning : Analyze current and future system resource needs and plan capacity accordingly to accommodate traffic and growth. Performance Optimization : Identify and address bottlenecks, performance issues, and resource inefficiencies to optimize system performance. Infrastructure as Code : Implement infrastructure as code (IaC) principles, using tools like Terraform or Ansible to define and manage infrastructure configurations. Scalability : Ensure that systems can scale horizontally or vertically to meet increased demand, adjusting resources as needed. Security : Address security concerns, vulnerabilities, and compliance requirements, and work to maintain a secure environment. Release Management : Collaborate with development teams to plan and execute releases and updates in a way that minimizes disruptions to production systems. On-Call Rotation : Participate in an on-call rotation to respond to incidents outside of regular working hours and to ensure continuous service availability. Documentation : Maintain comprehensive documentation for systems, configurations, and procedures, facilitating knowledge sharing and troubleshooting. Disaster Recovery : Plan and test disaster recovery and backup strategies to ensure data and service continuity in case of major failures or disasters.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
