Experience
4 - 8 yrs
Job Location
Chennai, India
Vacancy
1
Designation
Hadoop Administrator
Job Type
Not specified
Job Description
Role: Hadoop Admin
EXP 4 to 8 years
Location Chennai / Bangalore.
NP Immediate to 30 days.
Key Responsibilities of a Hadoop Administrator
- Cluster Deployment and Maintenance : Deploy and maintain Hadoop clusters, ensuring they are configured correctly and optimized for performance. This includes setting up configurations like Core-Site, HDFS-Site, YARN-Site, and MapRed-Site.
- Node Management : Add or remove nodes in the cluster using tools like Cloudera Manager, Ambari, or Ganglia. Ensure proper resource allocation and monitor node performance.
- High Availability Configuration : Configure NameNode high availability to prevent downtime and ensure the cluster is always operational.
- Monitoring and Troubleshooting : Use monitoring tools like Nagios or Ganglia to track cluster health, connectivity, and performance. Troubleshoot application errors and resolve issues to prevent recurrence.
- Backup and Recovery : Perform regular backups of Hadoop data and implement recovery mechanisms to safeguard against data loss.
- Capacity Planning : Estimate and plan for storage and processing capacity based on the data volume and growth trends.
- Log Management : Manage and review Hadoop log files to identify and resolve issues proactively.
- Security and Resource Management : Implement security measures, manage user access, and ensure proper resource allocation across the cluster.
- Collaboration with Teams : Work closely with database, network, BI, and application teams to ensure seamless integration and performance of big data applications.
- Design implement migration strategies for traditional systems on Azure (Lift and shift/Azure Migrate)
- Expertise in Hadoop architecture, design and development of Big Data platform including large clusters, Hadoop ecosystem components in a multi-node cluster environment
- Strong knowledge on Hadoop concepts and its Ecosystem components such as HDFS, MapReduce, YARN, Sqoop, Hive, Oozie. Knowledge of Cluster coordination services using Zookeeper.
- Expertise in Hadoop job schedulers such as Fair scheduler and Capacity scheduler
- Built custom Airflow sensors and operators to extract DAG run metadata, task execution details, and lineage information using Airflow REST API and secure API tokens;
- Created a metadata wrapper service that captures real-time Kafka events (Schema Registry changes, topic creation, DAG triggers). Created automated performance regression suites (JMeter + custom Spark scripts) and integrated them into CI/CD pipelines.
- Delivered performance tuning recommendations and compliance reports to senior stakeholders, ensuring the migrated platform met banking SLAs for latency, throughput, cost, and regulatory requirements
- Owned end-to-end performance and vulnerability remediation for the Global Data Platform migration from Cloudera CDP to cloud-native GDP, achieving zero performance regression and successful production sign-off.
- Non-Functional Requirements (NFR) and Production Certification Assessment (PCA) sign-off for Airflow and Apache Ranger components; conducted security hardening, vulnerability scanning, patch management, and compliance validation
