TymblHub

© 2026 TymblHub

Hadoop Administrator

Home Credit
Posted on
Home Credit  logo

Experience
7 - 12 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Hadoop Administrator
Job Type
Not specified

Job Description

Job Description :


Platform Operations & Administration

  • Administer and support Hadoop ecosystem clusters (e.g., HDFS, YARN, MapReduce) and related services to meet SLA/SLI targets for availability and performance.
  • Perform cluster provisioning, configuration, tuning, and day-to-day operational support (start/stop, health checks, service validation).
  • Manage multi-cluster environments (dev/test/prod), including standardization of configurations and operational procedures.

Installation, Upgrades & Lifecycle Management

  • Plan and execute platform upgrades, patching, and migrations with minimal downtime (including rolling upgrades where applicable).
  • Maintain platform compatibility matrix, validate integrations, and coordinate release rollouts with stakeholders.
  • Manage vendor distributions and tooling (e.g., Cloudera/HDP/Apache), including repository, parcels/packages, and version control.

Monitoring, Incident & Problem Management

  • Implement robust monitoring/alerting for platform and infrastructure (CPU, memory, disk, network, I/O, JVM, services, queues, HDFS capacity, GC, etc.).
  • Lead troubleshooting for incidents, perform root cause analysis (RCA), and drive corrective/preventive actions.
  • Own operational runbooks, on-call readiness, escalation processes, and continuous improvement of response/MTTR.

Performance Tuning & Optimization

  • Tune HDFS/YARN for workload performance and stability (scheduler/queue policies, resource limits, container sizing, node settings).
  • Optimize cluster performance through JVM tuning, service parameters, OS/kernel tuning, and workload analysis.
  • Identify and remediate bottlenecks (skew, small files, disk hotspots, network saturation, metadata pressure, etc.)

Capacity Planning & Reliability Engineering

  • Forecast growth and perform capacity planning for storage, compute, and service scaling.
  • Drive high availability and disaster recovery readiness (NameNode HA, ZK quorum health, backup/restore, replication, snapshots, DR drills).
  • Ensure data durability and operational resilience through proactive maintenance and reliability reviews.

We are 5 days work from office

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.