Experience
7 - 12 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Hadoop Administrator
Job Type
Not specified
Job Description
Job Description :
Platform Operations & Administration
- Administer and support Hadoop ecosystem clusters (e.g., HDFS, YARN, MapReduce) and related services to meet SLA/SLI targets for availability and performance.
- Perform cluster provisioning, configuration, tuning, and day-to-day operational support (start/stop, health checks, service validation).
- Manage multi-cluster environments (dev/test/prod), including standardization of configurations and operational procedures.
Installation, Upgrades & Lifecycle Management
- Plan and execute platform upgrades, patching, and migrations with minimal downtime (including rolling upgrades where applicable).
- Maintain platform compatibility matrix, validate integrations, and coordinate release rollouts with stakeholders.
- Manage vendor distributions and tooling (e.g., Cloudera/HDP/Apache), including repository, parcels/packages, and version control.
Monitoring, Incident & Problem Management
- Implement robust monitoring/alerting for platform and infrastructure (CPU, memory, disk, network, I/O, JVM, services, queues, HDFS capacity, GC, etc.).
- Lead troubleshooting for incidents, perform root cause analysis (RCA), and drive corrective/preventive actions.
- Own operational runbooks, on-call readiness, escalation processes, and continuous improvement of response/MTTR.
Performance Tuning & Optimization
- Tune HDFS/YARN for workload performance and stability (scheduler/queue policies, resource limits, container sizing, node settings).
- Optimize cluster performance through JVM tuning, service parameters, OS/kernel tuning, and workload analysis.
- Identify and remediate bottlenecks (skew, small files, disk hotspots, network saturation, metadata pressure, etc.)
Capacity Planning & Reliability Engineering
- Forecast growth and perform capacity planning for storage, compute, and service scaling.
- Drive high availability and disaster recovery readiness (NameNode HA, ZK quorum health, backup/restore, replication, snapshots, DR drills).
- Ensure data durability and operational resilience through proactive maintenance and reliability reviews.
We are 5 days work from office
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
