Experience
5 - 10 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Operations Engineer
Job Type
ONSITE
Job Description
We are seeking a skilled MLOps Platform Engineer to build and maintain a robust platform for the entire Machine Learning (ML) lifecycle. This role involves automating ML model training, testing, deployment, and monitoring, integrating seamlessly with our existing infrastructure to accelerate ML innovation.
;
Key responsibilities:
- CI/CD for ML : Develop and manage CI/CD pipelines for ML model code and infrastructure, covering unit, integration, and deployment to all environments.
- Automated ML Training : Design and implement pipelines for repeatable model training, automatic sweeps, data processing, and hyperparameter optimization.
- Include scheduling, queuing, and cost monitoring for training runs.
- Support training on cloud and on-premise GPU resources.
-
- Model Monitoring : Establish tools and dashboards for continuous monitoring of deployed ML models, ensuring health, availability, and performance.
- Platform Integration : Ensure the MLOps platform integrates with company workflows, data pipelines, and computing architecture.
- MLOps Best Practices : Promote and implement best practices for reproducibility, version control, and governance across the ML lifecycle.
Skills:
- Strong experience MLOps, DevOps, or ML Engineering, focusing on ML infrastructure
- Programming: Strong proficiency in Python and ML frameworks ( TensorFlow, PyTorch, Scikit-learn ).
- CI/CD Orchestration: Experience with CI/CD tools (eg, Jenkins, GitLab CI, GitHub Actions) and workflow orchestrators (eg, Airflow).
- Model Serving De
