Job Description
We are seeking an experienced DevSecOps Engineers with 5-8 years of hands-on experience to design, implement, and maintain secure, scalable, and highly available infrastructure with a strong focus on Cloud & AI/ML platforms. Will be a key contributor in bridging development, security, and operations, ensuring our Cloud/AI systems are resilient, performant, secure, and production-ready.
Key Skills:
Python, AWS, Claude AI or any other AI, Gitlab CI/ CD, Terraform
Key Responsibilities
Site Reliability Engineering
Design, build, and maintain highly available, scalable, and fault-tolerant distributed systems
Define and track SLIs, SLOs, and SLAs; drive reliability improvements based on error budgets
Lead incident response, conduct blameless post-mortems, and implement preventive measures
Build and improve observability stack (monitoring, logging, tracing, alerting)
Automate toil reduction through tooling and self-healing infrastructure
Perform capacity planning and optimize system performance and cost efficiency
Implement chaos engineering practices to proactively identify system weaknesses
Soft Skills
Strong analytical and problem-solving abilities
Excellent communication skills (written and verbal)
Ability to work under pressure during incidents
Proactive mindset with ownership mentality
Collaborative approach with cross-functional teams
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
