Job Description
Google Infra AI Devops Engineer
Position Overview
The Google Cloud AI DevOps Engineer (8+ years) is a highly skilled professional responsible for designing, building, and managing end-to-end AI/ML pipelines and infrastructure on Google Cloud Platform (GCP).
This role blends expertise in AI/ML development, cloud infrastructure, and DevOps best practices, enabling seamless deployment, continuous integration/continuous delivery (CI/CD), and monitoring of AI-driven solutions.
The engineer works closely with data scientists, AI developers, and infrastructure teams to deliver scalable, automated, and efficient AI products.
This role is designed for professionals who can bridge AI development with robust DevOps methodologies to enable scalable, automated, and reliable AI deployments on GCP, leveraging a blend of native cloud services and open-source technologies.
Qualifications & Certifications:
- Bachelor s degree in IT/ Engineering/MBA or other management qualification.
- GCP Devops / ML Engineer Architect Certification
- Kubernetes Certification [CKA CKAD etc.]
- Terraform Associate
GCP Services-Focused Responsibilities
- Design, implement, and maintain automated AI/ML pipelines leveraging GCP services such as Cloud Build, Cloud Functions, and Kubernetes Engine (GKE).
- Develop and manage infrastructure using Infrastructure as Code tools (Terraform, Cloud Build, Cloud Deploy) ensuring repeatable, scalable deployments.
- Develop custom Agents using ADK, crew AI, Langraph etc. Framework and deploy on runtime like Cloud Run, GKE, Agent Engine.
- Create, maintain, and optimize Dockerfiles and container images to package AI models and services for consistent deployment.
- Integrate open-source DevOps tools such as Jenkins, GitHub Actions, Bitbucket Pipelines, ArgoCD, and Tekton to support CI/CD workflows for AI/ML projects.
- Collaborate with AI/ML scientists to operationalize models from training to production, including version control, testing, and performance monitoring.
- Implement monitoring and logging solutions (Cloud Operations suite, Prometheus, Grafana) to ensure reliability, availability, and performance of AI applications.
- Automate the deployment of AI solutions with a focus on reproducibility, scalability, and security best practices on GCP.
- Manage source control repositories on GitHub, Bitbucket, or other platforms with branching strategies, code reviews, and release management.
- Ensure compliance with compliance, security policies, and data governance standards across AI/DevOps workflows.
- Provide technical guidance and documentation for AI DevOps processes, supporting cross-functional teams in adopting best practices.
Required Skills and Expertise
- Proven experience designing and managing AI/ML pipelines and infrastructure on Google Cloud Platform with tools like Vertex AI, GKE, Cloud Build, and Cloud Functions.
- Strong knowledge of DevOps practices and open-source CI/CD tools, including Jenkins, GitHub Actions, Bitbucket Pipelines, and Kubernetes-native solutions like ArgoCD and Tekton.
- Expertise in containerization technologies such as Docker, and experience writing Dockerfiles and managing container registries.
- Hands-on skills in Infrastructure as Code (Terraform, Deployment Manager) and automation scripting (Python, Bash).
- Familiarity with source control platforms (GitHub, Bitbucket) enabling branching, pull requests, and release management.
- Experience in monitoring and logging AI systems using Google Cloud Operations Suite, Prometheus, and Grafana.
- Strong collaboration skills working with data scientists, AI engineers, and infrastructure teams to deliver production-ready AI solutions.
- Understanding cloud security, identity, and access management related to AI/DevOps workflows.
- Background in troubleshooting, performance tuning, and scaling AI workloads on GCP.
- Relevant Google Cloud certifications such as Professional Cloud DevOps Engineer or AI Engineer certifications preferred.
