Job Description
About the Role:
We are looking for a highly skilled Infrastructure Engineer to design, build, and maintain scalable, reproducible lab environments to support integration with third-party vendors. This role is critical in enabling fast, reliable onboarding of external systems by providing production-like environments, automation, and debugging support for integration workflows. You will work closely with engineering, product, and partner teams to accelerate integration development, testing, and validation across diverse vendor ecosystems.
Basic Qualifications:
4+ years of experience in infrastructure engineering, DevOps, or similar roles with hands-on Python scripting capabilities.
Hands-on experience with infrastructure-as-code tooling Terraform for VM provisioning, Packer for VM templating, and Ansible for configuration management, including writing and maintaining roles, modules, and playbooks in a team setting with PR-based review
Practical experience deploying and managing workloads on at least two of: VMware vSphere/ESXi, AWS, Azure, GCP, or Proxmox, with a solid understanding of VM lifecycle, networking, and snapshots
Working knowledge of containerization and orchestration (Docker, Kubernetes, or OpenShift) is a plus
Linux system administration skills, including shell scripting (and/or Python), for automating provisioning and configuration tasks
Comfort building and consuming REST APIs, and working in a platform where provisioning is driven programmatically rather than through manual console steps
Familiarity with CI/CD pipelines and version control (Git/GitHub), including pull-request-based contribution and peer review workflows
A collaborative engineering mindset writing environment definitions, playbooks, or modules that other engineers across the organization will read, reuse, and build on
Strong troubleshooting and root-cause analysis skills, with experience using monitoring/observability tooling (e.g., Splunk, Grafana, Nagios) to validate environment health and performance.
Nice to have:
Experience supporting enterprise-scale virtualization deployments (e.g., VMware Horizon, App Volumes, vSphere) across regulated or large client environments.
Enough security domain knowledge to build environments that are meaningful for detection engineering and SOC workflows exposure to attack scenarios, incident response, or threat-informed defence is valued
Experience with Windows Active Directory environments designing, deploying, or troubleshooting AD, DNS, DHCP, and Windows Server infrastructure is a strong plus.
Background in incident and change management practices, including pre-production review steps to minimize downtime
Cloud certifications (e.g., AWS SysOps, Terraform Associate) or virtualization certifications (e.g., VCP/VCAP/VCIX)
Key Responsibilities
Lab Environment Design Management
Design and build isolated, secure, and reusable lab environments (cloud, on-prem, or hybrid)
Create multi-component setups including controllers, agents, databases, and endpoint systems
Support a variety of vendor deployment models (SaaS, VM-based, containerized)
Ensure lab environments are representative of production scenarios where needed
Automation Infrastructure as Code
Build and manage environments using Infrastructure as Code (Terraform, ARM/Bicep, CloudFormation)
Develop automated provisioning pipelines for repeatable lab setups
Maintain standardized templates and golden images
Version control infrastructure and configurations using Git-based workflows
Ability to build and consume REST APIs, work within a platform where all provisioning is driven programmatically, and contribute to a system where the UI and CLI are clients on top of a shared API.
Networking Connectivity
Configure and troubleshoot:
Virtual networks, subnets, routing
Load balancers, NAT gateways, VPNs
Firewall and proxy configurations
Simulate enterprise scenarios such as:
Restricted/segmented networks
Multi-region or cross-tenant deployments
Integration Enablement
Support integration workflows involving:
REST/GraphQL APIs
Webhooks and event-driven systems (Kafka, Event Hub, MQTT)
Enable testing of:
Authentication (OAuth2, API keys, certificates)
Failure scenarios, retries, and rate limits
Build mock services and simulators for vendor dependencies
Observability Troubleshooting
Implement logging, monitoring, and tracing (e.g., ELK, Grafana, Azure Monitor)
Provide debugging support through:
Traffic inspection tools
API tracing and log correlation
Assist teams in diagnosing and resolving integration issues efficiently
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
