Job Description
Job Title: Site Reliability Engineer (SRE) For AI Platform/Application
Key Skills
1. Cloud and Infrastructure: Azure, on-prem systems, good-to-have GCP.
2. Kubernetes: Deploy, manage, and troubleshoot services.
3. Observability: ELK/EFK, Grafana, Prometheus, Dynatrace, strong log analysis; Splunk is a plus.
4. Automation and Scripting: Bash and Python.
5. Programming: Python(preferred), Java
6. AI and Agentic Systems: LangChain/CrewAI/AutoGen (anyone), LangGraph, RAG, LLM basics, evals for AI observability, NLP/deep learning/transformer fundamentals and Prompt Engineering Skills.
7. SRE Core: SLA, SLI/SLO, alert configuration and tuning, production troubleshooting, end-to-end application flow understanding, ServiceNow incident handling.
8. Reliability and Platform: Data center failover and networking debugging knowledge.
9. DevOps: CI/CD pipelines, GitHub Actions, GitHub workflows.
10. Databases: Vector Database, PostgreSQL and Couchbase.
11. Practical AI: Should have built at least one real-time agentic workflow(Highly Preferred).
Primary Responsibilities
1. Ensure reliability, availability, and performance of production systems.
2. Monitor, triage, and resolve incidents and alerts.
3. Automate operational tasks and reduce manual toil.
4. Drive observability and RCA improvements.
5. Partner with engineering teams to integrate AI capabilities into existing architecture.
Regards,
Lenin ND
lenin.nd@atos.ai
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
