Experience
5 - 10 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Infrastructure Support Specialist
Job Type
ONSITE
Job Description
Project Role : Infra Tech Support Practitioner
Project Role Description : Provide L1L2 technical support for infrastructure, systems, and applications in production and development environments, both remotely and onsite, following defined operating models and processes. Act as the primary interface with users/clients to accurately diagnose issues and deliver effective resolutions across infrastructure, operating systems, networks, and application platforms, ensuring service stability and quality support.
Must have skills : Site Reliability Engineering
Good to have skills : NA
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
As an Infra Tech Support Practitioner, a typical day involves delivering first and second level technical assistance for infrastructure, systems, and applications across both production and development settings. This role requires working closely with users and clients to identify and resolve issues efficiently, whether remotely or onsite. The practitioner follows established operational procedures to maintain service reliability and quality, ensuring smooth functioning of various technology platforms. Collaboration with different teams and adherence to defined processes are key aspects of the daily routine, aimed at sustaining optimal infrastructure performance and user satisfaction.
Roles Responsibilities:
1. Infrastructure as Code Automation
Automate Everything: Eliminate manual interventions by designing and implementing end-to-end infrastructure automation using tools like Terraform or Ansible.
CI/CD Pipeline Management: Build, optimize, and maintain robust CI/CD pipelines to ensure seamless, automated software delivery and infrastructure deployments.
Configuration Management: Define and manage system configurations via code to maintain consistency across Development, Staging, and Production environments.
2. Reliability, Scalability Performance
Architecture Design: Partner with development teams to architect scalable, fault-tolerant, and high-performance cloud infrastructure.
Capacity Planning: Proactively monitor system resource utilization and plan capacity upgrades to handle traffic spikes and growth.
Chaos Engineering: Introduce controlled failures into systems to test resilience and proactively identify architectural weaknesses.
3. Monitoring, Observability Incident Response
Observability Stack: Design and implement comprehensive monitoring, logging, and alerting systems (e.g., Newrelic ) to establish deep system visibility.
On-Call Incident Management: Participate in an on-call rotation. Lead the triage and resolution of critical production incidents.
Blameless Post-Mortems: Conduct thorough root-cause analysis (RCA) for infrastructure failures and formulate actionable plans to prevent recurrence.
Professional Technical Skills:
Must To Have Skills:
Category Required Skills Tools
Cloud Platforms:- Advanced expertise in AWS (Cloud networking, IAM, security, compute, and storage services).
Containers Orchestration:- Deep operational knowledge of Docker and Kubernetes (EKS, GKE, or AKS), including Helm charts and service meshes.
Infrastructure as Code (IaC) Strong proficiency in Terraform or CloudFormation.
Programming Scripting:- Proficiency in Python, or Bash for building internal tooling and automation scripts.
CI/CD Tools Hands-on experience with GitHub Actions, GitLab CI, Jenkins.
Observability Experience configuring in New Relic.
OS Networking Strong Linux administration skills and a solid understanding of TCP/IP, DNS, Load Balancing, and HTTP/S.
Qualification 15 years full time education
