Job Description
Key Responsibilities
Infrastructure Reliability & Operations
- Manage and maintain large-scale hybrid infrastructure environments across on-premises and cloud platforms.
- Ensure high availability, resiliency, scalability, and performance of critical business applications and infrastructure services.
- Lead proactive monitoring, incident response, problem management, and infrastructure health assessments.
- Perform root cause analysis (RCA) for major incidents and implement preventative measures.
- Drive service availability improvements through automation and operational excellence initiatives.
- String understanding on Application Modernization , Replication, Failover Strategy
Cloud & Hybrid Infrastructure
- Support and optimize Azure, AWS, Azure Local (Azure Stack HCI), and hybrid cloud environments.
- Participate in cloud migration, modernization, and transformation initiatives.
- Design and implement resilient infrastructure architectures supporting business continuity requirements.
- Collaborate with application, cloud, security, and infrastructure teams during cloud adoption programs.
Disaster Recovery & Business Continuity
- Lead DC/DR planning, execution, testing, and governance activities.
- Define and validate RPO/RTO objectives.
- Execute disaster recovery drills, failover testing, and recovery validation exercises.
- Ensure operational readiness for business continuity and infrastructure resiliency requirements.
Automation & Infrastructure as Code (IaC)
- Implement infrastructure provisioning and configuration management using Terraform and Ansible.
- Develop and maintain automation workflows to improve operational efficiency.
- Automate routine operational activities, infrastructure deployments, compliance checks, and service restoration tasks.
- Support Infrastructure-as-Code standards, reusable templates, and automation governance.
SRE & Reliability Engineering
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Establish reliability metrics and operational KPIs.
- Support reliability-focused architecture reviews and capacity planning activities.
- Drive continuous service improvement initiatives through data-driven operational insights.
Observability & Monitoring
- Build and enhance monitoring, alerting, and observability dashboards.
- Establish predictive monitoring and proactive issue detection capabilities.
- Analyze infrastructure and application telemetry to identify performance bottlenecks.
- Improve visibility using enterprise monitoring and observability platforms.
Platform & Virtualization Management
- Administer and optimize VMware infrastructure.
- Manage SAN/Storage platforms, storage performance, and capacity planning.
- Support backup and recovery solutions, ensuring recovery readiness and compliance requirements.
- Maintain secure, scalable, and governed infrastructure environments.
Governance & Compliance
- Ensure adherence to enterprise governance, operational standards, security policies, and regulatory requirements.
- Participate in architecture reviews, risk assessments, and operational readiness reviews.
- Support audit requirements, infrastructure documentation, and change management processes.
Required Skills & Experience
Infrastructure Technologies
- Linux Administration (RHEL, CentOS, Ubuntu, SUSE)
- VMware vSphere, ESXi, vCenter
- SAN/Storage Administration
- Backup & Recovery Solutions
- Data Center Operations
- High Availability (HA) and Disaster Recovery (DR)
- Integration and Middelware
Cloud Platforms
- Microsoft Azure
- Amazon Web Services (AWS)
- Azure Local / Azure Stack HCI
- Cloud Migration and Modernization
Automation & IaC
- Terraform
- Ansible
- Infrastructure Automation
- Configuration Management
Monitoring & Observability
- Monitoring Platforms
- Logging & Alerting Solutions
- Performance Management
- Dashboard Development
- Operational Analytics
Scripting & Automation
- Python (Preferred)
- Shell Scripting (Bash)
- REST APIs & Automation Integrations
Preferred candidate profile
Education : Bachelor's degree in Computer Science, Engineering, MCA
Primary Stack : On-Premise Infrastructure, Middleware, Database, Application Hosting, Load Balancing, DNS
Infrastructure Technologies- Linux Administration (RHEL, CentOS, Ubuntu, SUSE), VMware vSphere, ESXi, vCenter SAN/Storage Administration, Disaster Recovery, Back Up, Storage Replication
Monitoring & Observability
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
