Job Description
Company Overview
Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model". With 25+ years of transforming enterprises, weve evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.
With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps - delivered through strategic partnerships with leading hyperscalers.
Were building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence - RevolutionAIzing Transformations.
ABOUT THIS ROLE
You are the execution engine behind the AI Platform organization. You take architecture designs and platform standards defined by the Principal AI Architect and turn them into production-ready, repeatable, and secure deployment systems. Your mandate is simple: reduce deployment friction, accelerate model releases, and ensure AI infrastructure can scale reliably across cloud and on-premises environments.
This is not traditional DevOps. This is not infrastructure administration. This is hands-on AI Platform Engineering focused on deployment automation, GPU infrastructure, CI/CD for AI workloads, and DevSecOps at scale.
You will build the pipelines that ship AI models to production, automate infrastructure provisioning across multi-cloud environments, manage GPU-enabled Kubernetes clusters, and implement the controls that ensure every deployment meets reliability and security standards.
WHAT YOU WILL DO
AI Platform Deployment Engineering
- Build and maintain CI/CD pipelines for LLM and ML model releases
- Implement model versioning, staged rollout, automated rollback, and release governance workflows
- Design deployment automation that reduces release times from hours to minutes
- Standardize deployment processes across cloud and on-premises AI environments
- Ensure deployment reliability, consistency, and repeatability at scale
Infrastructure Automation & Provisioning
- Develop infrastructure-as-code using Terraform or Pulumi across AWS, Azure, GCP, and OpenShift environments
- Automate provisioning of GPU and CPU infrastructure for training and inference workloads
- Build self-service deployment capabilities that allow rapid environment creation
- Standardize cloud infrastructure patterns and reusable deployment modules
- Automate resource lifecycle management, scaling, and environment configuration
GPU Platform Operations
- Provision and manage GPU nodes within Kubernetes environments
- Configure NVIDIA Device Plugin, node selectors, taints, tolerations, and GPU resource allocation strategies
- Maintain CUDA drivers, GPU dependencies, and cluster-level GPU configurations
- Optimize GPU utilization across multiple AI workloads and deployment environments
- Support scalable AI inference and training infrastructure requirements
DevSecOps & Governance
- Implement security gates across deployment pipelines
- Integrate container image scanning, infrastructure-as-code scanning, secrets detection, and policy enforcement
- Build automated compliance checks into AI deployment workflows
- Ensure security requirements are validated before production releases
- Maintain deployment standards across multi-cloud and on-premises environments
Observability & Reliability Engineering
- Implement monitoring and alerting for AI serving infrastructure
- Track model latency, throughput, system health, error rates, and GPU utilization
- Build dashboards and operational visibility for platform health
- Implement blue-green deployment strategies and automated rollback mechanisms
- Improve platform reliability, availability, and operational efficiency
MUST HAVE
- Built CI/CD pipelines that successfully shipped AI or ML models to production environments
- Hands-on experience with Terraform or Pulumi at production scale across at least two cloud providers
- Provisioned and managed GPU-enabled Kubernetes environments including NVIDIA Device Plugin, node selectors, and resource management
- Practical DevSecOps experience with tools such as Trivy, Checkov, Vault, OPA, or equivalent technologies
- Strong Linux automation and scripting expertise using Bash and/or Python
- Hands-on experience deploying and managing Red Hat OpenShift operators in on-premises environments
- Minimum 4+ years of DevOps, Platform Engineering, or Infrastructure Engineering experience
- Minimum 2+ years supporting AI, ML, GenAI, or AI Platform environments
GOOD TO HAVE
- NVIDIA MIG partitioning for multi-tenant AI serving environments
- Helm chart development and optimization for ML serving workloads
- GitOps platforms such as ArgoCD or Flux
- GPU infrastructure cost optimization including spot instances, reserved capacity, and rightsizing strategies
- Kubernetes platform administration at enterprise scale
- Experience supporting large-scale inference platforms and model serving systems
- Familiarity with AI platform frameworks and MLOps tooling ecosystems
Why Choose Trianz
- AI Platform at Scale: Build the infrastructure that powers enterprise-grade AI platforms, model deployments, and next-generation AI services across global clients.
- AI-First Future: Be part of a company where AI sits at the center of product innovation, customer transformation, and platform strategy.
- Global Impact: Support AI initiatives across industries, geographies, and some of the worlds leading enterprise organizations.
- Technical Ownership: Own critical deployment systems, automation frameworks, and platform capabilities that directly impact AI delivery velocity.
- High-Growth Environment: Work in a fast-paced culture that rewards engineering excellence, innovation, and proactive problem-solving.
- Entrepreneurial Spirit: Enjoy the freedom to build, automate, and optimize without unnecessary bureaucracy while delivering meaningful business outcomes at global scale.
