Lead SRE

Techverito Software Solutions LLP
Posted on

Experience
4 - 8 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Site Reliability Engineer Lead
Job Type
Not specified

Job Description

Job Summary

We are looking for a highly skilled Lead SRE / Infrastructure DevOps Engineer with strong expertise in Microsoft Azure, Infrastructure as Code (IaC), DevOps automation, and Oracle Database Platform . The ideal candidate will be responsible for designing, implementing, and managing secure, scalable, and highly available cloud infrastructure while driving automation, operational excellence, and platform reliability. This role requires hands-on experience in Azure cloud services, CI/CD, infrastructure automation, Oracle production environments, and Site Reliability Engineering (SRE) practices.


Job Requirements

Infrastructure & Cloud
  • Strong hands-on experience with Microsoft Azure infrastructure and cloud-native services.

  • Experience managing enterprise infrastructure including networking, compute, storage, identity, and platform services.

  • Strong understanding of Azure governance, subscriptions, resource groups, policies, and management groups.

  • Experience with Azure services including Virtual Machines, Virtual Networks, AKS, App Services, Function Apps, Storage Accounts, Azure SQL, and Azure Monitor.

DevOps & Automation
  • Strong experience with Infrastructure as Code using Terraform, ARM Templates, or Bicep.

  • Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions.

  • Strong scripting skills using PowerShell, Python, Bash, or Go.

  • Experience with Git, automation frameworks, and Infrastructure as Code.

  • Good understanding of container platforms including Docker and Kubernetes (AKS preferred).

Oracle Platform
  • Good understanding of Oracle Database architecture and administration concepts.

  • Experience supporting Oracle Database infrastructure in production environments.

  • Experience with Oracle performance tuning using AWR, ADDM, ASH, OEM, execution plans, and SQL optimization.

  • Familiarity with Oracle RMAN, Data Guard, backup and recovery, and high availability concepts.

  • Experience working closely with Oracle DBAs during production troubleshooting and performance optimization.

Security & Monitoring
  • Experience implementing Azure RBAC, Managed Identities, Key Vault, Private Endpoints, Azure Policy, and Defender for Cloud.

  • Strong understanding of cloud security, governance, compliance, and enterprise security best practices.

  • Experience with Azure Monitor, Log Analytics, Application Insights, Microsoft Sentinel, Prometheus, Grafana, or similar monitoring platforms.

  • Strong understanding of SRE principles including observability, reliability engineering, automation, and incident management.

Job Responsibilities

  • Design, deploy, and manage enterprise cloud infrastructure on Microsoft Azure.

  • Build and maintain Infrastructure as Code (IaC) using Terraform, ARM Templates, or Bicep.

  • Develop and maintain CI/CD pipelines using Azure DevOps and GitHub Actions.

  • Administer Azure services including Virtual Machines, Virtual Networks, AKS, App Services, Function Apps, Storage Accounts, Azure SQL, and Azure Monitor.

  • Implement platform automation and configuration management using PowerShell, Python, Bash, or Go.

  • Design and maintain secure Azure environments using RBAC, Managed Identities, Key Vault, Private Endpoints, Azure Policies, and governance frameworks.

  • Monitor infrastructure health, performance, availability, and security using Azure Monitor, Log Analytics, Application Insights, Microsoft Sentinel, and other enterprise monitoring tools.

  • Manage Oracle Database platform operations including installation, patching, backup validation, disaster recovery, storage optimization, and capacity planning.

  • Collaborate with Oracle DBAs to resolve infrastructure and database performance bottlenecks.

  • Analyze Oracle performance reports (AWR, ASH, execution plans) and support infrastructure optimization.

  • Troubleshoot production incidents across cloud infrastructure, middleware, and Oracle platforms while driving root cause analysis and preventive improvements.

  • Work closely with Architecture, Security, Application, and Operations teams to deliver secure, scalable, and resilient platform solutions.

  • Prepare technical documentation, operational runbooks, architecture diagrams, and standard operating procedures.

  • Drive automation initiatives that improve operational efficiency, platform reliability, and service availability.


Benefits

  • Innovative Engineering: Collaborative, fail-fast, flat hierarchy. Fosters learning, initiative, curiosity.

  • Masterful Development: Emphasizes clean code, SOLID principles, TDD/BDD. Utilizes robust CI/CD and polyglot engineering.

  • Continuous Growth: Structured mentorship, masterclasses, Geeknights, workshops, continuous skill enhancement, blog contributions.

  • Agile & Client-Centric: Adopts Agile (Scrum, XP), promotes project ownership and deep client understanding for impactful solutions.

  • Supportive Environment: Healthy work-life balance, flexible schedules, comprehensive benefits (generous leave), strong team-building.



Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.