Job Description
Position: MLOps Engineer
Level: Senior
Experience: 8 10 years
Location: Noida (hybrid)
Budget: 16-18 LPA
About the role
You own the infrastructure that keeps our AI products running and running accurately long after launch. This is a combined DevOps + MLOps role because the two disciplines now overlap heavily in production AI systems, and wed rather have engineers who think across both than two silos that hand work over a wall.
You build the CI/CD pipelines, manage the cloud infrastructure, and operate the model lifecycle: deployment, monitoring, drift detection, retraining, and rollback. Our commercial model includes keeping client AI systems from silently degrading the 10 15% that unmonitored chatbots typically drop within months of launch.
This is a production-first role. If something breaks in a BFSI clients loan workflow at 2am, youre the person who finds out before they do.
Required qualifications
- 8 10 years in DevOps, SRE, or platform engineering, with the last 2+ years operating production ML or LLM systems (not just shipping models from notebooks to a REST endpoint).
- Strong Python (for tooling and pipeline code) and strong Bash. Working knowledge of Go or one statically-typed language is a plus.
- Production experience with at least one managed ML platform AWS SageMaker, Azure ML, or Google Vertex AI including endpoints, pipelines, and model registries.
- Production experience with Kubernetes and Terraform (or equivalent IaC).
- Hands-on experience with at least one CI/CD platform GitHub Actions, GitLab CI, Jenkins, CircleCI, Argo CD.
- Demonstrable experience with observability tooling Prometheus/Grafana, Datadog, CloudWatch, or the equivalent on Azure/GCP.
- Hands-on experience with drift detection and model monitoring Evidently, WhyLabs, Arize, Fiddler, or a custom-built equivalent you can speak to in detail.
- Solid grasp of cloud cost optimisation youve actually cut cloud bills, not just talked about it.
Preferred qualifications
- LLM-specific observability experience LangSmith, LangFuse, Helicone, Arize Phoenix, or equivalent.
- Experience deploying and operating open-source LLMs on private infrastructure Llama 3, Mistral, vLLM, TGI, Ollama, on bare metal or VPC.
- GPU infrastructure experience provisioning, scheduling, cost management (A100/H100/L4 fleets).
- Experience with feature stores Feast, Tecton, SageMaker Feature Store.
- Familiarity with DPDP Act 2023, HIPAA, and sector-specific compliance requirements (RBI cybersecurity guidelines, IRDAI, etc.).
- Prior on-call experience for production systems with paying clients.
- Experience supporting consulting or services-led engagement models.