Job Description
Site Reliability Engineer (SRE)
CORE TECHNICAL REQUIREMENTS SPECIFICATION
Position Summary Focus Area: Infrastructure Stability, Cloud Orchestration, and Automation
We are seeking an experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will bridge the gap between application deployment and robust infrastructure operations, utilizing a software defined approach to system reliability, visibility, and performance scaling.
Core Technical Requirements
Containerization & Orchestration
Strong conceptual and operational knowledge of Kubernetes (K8s) architecture. Candidate must understand control plane mechanics, cluster networking, workload scheduling, and troubleshooting paradigms inside out.
CI/CD & Progressive Delivery
Practical experience with application deployment workflows and progressive delivery strategies (e.g., Canary rollouts, Blue/Green patterns). Ability to confidently diagnose and fix deployment pipeline friction or structural lifecycle failures.
Cloud Infrastructure Platforms
Proven working experience within a major public cloud provider environment (such as AWS, GCP, or Azure). Ability to navigate native services, resource provisioning frameworks, and identity management confidently.
Automation & Scripting
Proficiency in at least one modern automation or system scripting language (e.g., Bash, Python, or Go). Demonstrable ability to write clean, reusable code to minimize administrative toil and build reliable automation tools.
Core System Networking
Foundational understanding of network infrastructure, routing principles, load balancing strategies, and core protocols (TCP/IP, DNS, HTTP). Ability to accurately diagnose complex interconnectivity or service communication bugs.
