Job Description
Current Tech Stack & Infrastructure Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL) Orchestration: Kubernetes on EKS with Karpenter for node management Infrastructure as Code: Terraform (multiple repositories, experiencing drift and collaboration challenges) Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and runbooks CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm deployments, Rundeck for scripted operations GitOps: ArgoCD for Kubernetes workloads Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL Aurora, and Pulsar for event streaming APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts) Additional Monitoring: Started Prometheus clusters within EKS for more verbose metrics alongside Datadog Scale & Workload Handling hundreds of thousands of events per second via API (mix of synchronous and asynchronous processing) EC2 instances: M5/M6 xlarge instances, scaling between 15-100 instances per environment depending on load Currently transitioning workloads from EC2 to Kubernetes (K8s workload still smaller than core application) Real-time workloads requiring minimal downtime (minutes not hours for maintenance windows) Key Operational Challenges Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation and multi-person collaboration issues Legacy Environments: Some environments set up entirely manually, never in Terraform, with major operational challenges to migrate while keeping them online Terraform Migration: Proven migration patterns in test environments, but moving to production challenging due to real-time workload requirements Alert Management: Receiving too many alerts, need prioritization and structured approach to reduce noise and recategorize/adjust thresholds Alert Distribution: Historically all alerts went to one tech ops team instead of being distributed to five different dev teams; working to shift left closer to developers Service Catalog: Still defining service catalog and ownership model Runbooks: Exist in Datadog but need more polish and structure Database Scaling: Operational challenges with database scaling being addressed through data store segmentation Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to cost Manual Changes: Infrastructure changes often implemented manually first, then imported to Terraform and rolled out to other environments
Top 5 Required Skill Sets
- Strong Terraform and CI/CD process expertise
- Experience breaking down alerts to identify critical vs. non-critical and handling frequent alarms (Datadog-specific experience preferred)
- SLA/SLI definition experience (team currently being asked to do this without prior experience)
- AWS infrastructure expertise, DevOps-heavy background
- Monitoring and alerting experience (Datadog preferred, though Prometheus/Grafana experience also valuable)
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
