Job Description
What You ll Be Doing:
Managing and optimizing large-scale AWS cloud infrastructure
Driving Kubernetes (EKS) adoption and workload migration
Automating infrastructure using Terraform and GitOps practices
Enhancing platform reliability, observability, and operational excellence
Defining and implementing SLI/SLO frameworks
Improving monitoring, alerting, and incident response processes
Collaborating with development teams to build highly available and resilient systems
Key Skills Required:
AWS (EKS, EC2, RDS Aurora, ElastiCache, Control Tower)
Kubernetes (Multi-Cluster & Multi-Environment)
Terraform (Infrastructure as Code)
ArgoCD, Atlantis, GitOps
Karpenter & KEDA
Datadog, Prometheus, Monitoring & Alerting
CI/CD Pipelines
SLI/SLO Implementation and Site Reliability Engineering Practices
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
