Job Description
Position: Staff Systems Engineer Cloud Infrastructure
Grade: IT4
Location: Noida/Bangalore
Job Summary
We are seeking a Staff Systems Engineer Cloud Infrastructure to design, implement, and manage large-scale public, private, and hybrid cloud platforms , with deep expertise in OpenStack and cloud IaaS technologies . This role will serve as a technical authority and architect for cloud infrastructure platforms that power mission-critical workloads across EDA, Systems design, AI/ML, and high-performance computing environments.
The ideal candidate brings hands-on engineering depth , strong systems thinking , and a proven track record of leading complex cloud platform initiatives from architecture through production operations while mentoring senior engineers and influencing platform strategy across the organization.
Key Responsibilities
Cloud Platform Architecture & Strategy
-
Define and implement reference architectures for private cloud (OpenStack) and hybrid/multi-cloud environments.
-
Contribute towards architectural decisions for compute, storage, networking, and control plane services at scale.
-
Ensure platform consistency, scalability, resiliency, and operational excellence across regions and data centers.
-
Evaluate and guide adoption of emerging cloud and IaaS technologies aligned with business and platform strategy.
Private Cloud & OpenStack Leadership
-
Serve as a technical expert for OpenStack (Nova, Neutron, Cinder, Keystone, Glance, Heat, Octavia, etc. ).
-
Monitor, maintain, and optimize OpenStack deployments for performance, availability, upgradeability, and lifecycle management .
-
Lead complex upgrades, migrations, and capacity expansions with minimal customer impact.
-
Define best practices for multi-tenant isolation, quota management, and SLA-driven design .
Public Cloud & Hybrid Cloud Integration
-
Design and integrate workloads across AWS, Azure, and/or GCP alongside private cloud platforms.
-
Establish patterns for hybrid networking, identity federation, workload portability .
-
Provide technical guidance on cloud-native and IaaS-based workload placement decisions.
Infrastructure Engineering & Automation
-
Develop infrastructure-as-code and automation using tools such as Terraform, Ansible, Helm, and CI/CD pipelines.
-
Promote platform standardization, self-service provisioning, and repeatable operational workflows.
-
Perform SRE and operations activities to improve observability, reliability, and incident response .
Performance, Scale & Reliability
-
Perform root cause analysis for complex platform issues across compute, network, and storage layers.
-
Optimize cloud platforms for high-performance workloads , including AI/ML, batch, and HPC-style systems.
-
Monitor and improve key reliability, performance, and capacity metrics.
Technical Leadership & Collaboration
-
Act as a senior technical authority across cloud, infrastructure teams.
-
Mentor senior engineers; raise the technical bar across the organization.
-
Influence roadmap priorities through deep technical insight and clear communication with engineering and technical leaders.
-
Partner closely with security, networking, hardware, and application teams to deliver integrated solutions.
Required Qualifications
Experience
-
8+ years of experience in cloud infrastructure, distributed systems, or large-scale platform engineering.
-
Deep, hands-on experience with OpenStack in production environments.
-
Proven experience designing and operating large-scale private and hybrid cloud platforms .
-
Demonstrated leadership on complex, cross-team infrastructure initiatives.
Technical Expertise
-
Strong knowledge of:
-
OpenStack services (Nova, Neutron, Cinder, Keystone, etc. )
-
Linux systems, virtualization (KVM), and container platforms (Kubernetes)
-
Software-defined networking (SDN), L2/L3 networking, load balancing
-
Distributed storage systems (Ceph, object/block storage)
-
-
Experience with:
-
Public cloud IaaS (AWS, Azure, GCP)
-
Infrastructure automation and IaC
-
CI/CD pipelines supporting platform engineering
-
Preferred Qualifications
-
Experience supporting HPC/EDA workloads on cloud platforms.
-
Exposure to bare-metal provisioning , hardware lifecycle management, and data center operations.
-
Familiarity with SRE practices , reliability engineering, and large-scale observability platforms.
-
Strong written and verbal communication skills, with the ability to influence at senior and executive levels.
Educational Qualification:
BE/BTech/ME/MTech in CS/Electronics or equivalent
