Experience
6 - 11 yrs
Salary (CTC)
₹6.5L - ₹7.3L
Job Location
Bengaluru, India
Vacancy
3
Designation
Site Reliability Engineer
Job Type
Not specified
Job Description
Key Responsibilities
- Manage and support infrastructure powering AI/ML and Generative AI applications.
- Design and implement scalable, highly available, and secure platform solutions.
- Build automation to reduce operational toil and improve platform reliability.
- Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Operate Kubernetes-based environments, container platforms, and cloud services.
- Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting.
- Lead incident response, root cause analysis (RCA), and reliability improvement initiatives.
- Perform capacity planning, performance optimization, and cost management.
- Support GPU-based compute environments, AI model serving, and data pipelines.
- Implement disaster recovery, backup, resiliency, and security controls.
- Collaborate with engineering and security teams to deploy and operate AI services safely.
- Maintain operational documentation, runbooks, and best practices.
- Participate in on-call support and production incident management.
Required Skills & Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering.
- Strong programming/scripting skills in Python, Go, Java, or similar languages.
- Hands-on experience with Docker, Kubernetes, and container orchestration.
- Experience with AWS, Azure, or Google Cloud Platform (GCP).
- Strong knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
- Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry.
- Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing.
- Experience in incident management, RCA, performance tuning, and system scaling.
- Strong understanding of security, compliance, governance, and production support.
- Excellent communication and cross-functional collaboration skills.
Preferred Skills
- Experience supporting Generative AI, LLM, MLOps, ModelOps, or AI Platforms.
- Knowledge of GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling.
- Experience with Kafka, Spark, Flink, and distributed data processing frameworks.
- Familiarity with databases such as Snowflake, Redis, SQL, PostgreSQL.
- Understanding of Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving.
- Experience with Canary Deployments, Blue-Green Deployments, Chaos Engineering.
- Exposure to financial services or highly regulated environments.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
