Job Description
We are looking for a highly skilled Hadoop Administrator (SRE) to manage, support, and optimize large-scale distributed data platforms. The role focuses on ensuring high availability, reliability, scalability, and performance of Hadoop and Spark ecosystems running on Kubernetes and AWS infrastructure.
Key Responsibilities-
Administer, monitor, and support Hadoop ecosystems (HDFS, YARN, Hive, HBase, Spark, Kafka, etc.).
-
Manage and maintain Spark workloads for batch and streaming applications.
-
Deploy, manage, and troubleshoot Hadoop/Spark clusters on Kubernetes.
-
Ensure site reliability through proactive monitoring, alerting, incident response, and root cause analysis.
-
Automate infrastructure provisioning, configuration, and deployments using IaC and scripting.
-
Monitor system performance, capacity planning, and optimize cluster resource utilization.
-
Manage AWS infrastructure including EC2, S3, EKS, IAM, VPC, and related services.
-
Implement security best practices (Kerberos, Ranger, IAM, encryption, access control).
-
Perform cluster upgrades, patching, backup, and disaster recovery activities.
-
Collaborate with application, data engineering, and DevOps teams for seamless operations.
-
Provide L2/L3 production support and participate in on-call rotations.
-
5+ years of experience in Hadoop Administration / SRE / Big Data Operations.
-
Strong hands-on experience with:
-
Hadoop ecosystem (HDFS, YARN, Hive, HBase, Spark)
-
Apache Spark (performance tuning and troubleshooting)
-
Kubernetes (EKS preferred)
-
AWS Cloud services
-
-
Experience with monitoring tools (Prometheus, Grafana, CloudWatch, Ambari, Cloudera Manager).
-
Proficiency in Linux/Unix system administration.
-
Strong scripting skills in Shell, Python, or similar.
-
Experience with CI/CD pipelines and automation tools (Ansible, Terraform, Jenkins).
-
Solid understanding of SRE principles: SLAs, SLOs, SLIs, incident management.
