Job Description
Role & responsibilities
• Design, develop, and maintain scalable batch and streaming data pipelines using Spark and Scala.
• Build and optimize data processing workflows leveraging Hive for querying, transformations, and data warehousing needs.
• Implement reliable ingestion and integration patterns for high-volume datasets, ensuring data quality, consistency, and completeness.
• Develop reusable Spark jobs, libraries, and frameworks to standardize data engineering practices across teams.
• Tune Spark applications for performance (partitioning, caching, shuffles, memory management) and improve runtime efficiency.
• Work with stakeholders to understand data requirements and deliver well-modeled datasets for downstream consumption.
• Implement monitoring, alerting, and operational runbooks to ensure pipeline reliability and faster incident resolution.
• Perform code reviews, enforce engineering best practices, and contribute to continuous improvement of data platform standards.
Preferred Qualifications:
• Hands-on experience with Kafka for building streaming ingestion and event-driven data pipelines.
• Experience designing end-to-end data architectures (ingestion, processing, storage, and serving layers) for large-scale systems.
• Strong understanding of data partitioning strategies, file formats, and efficient processing patterns for big data workloads.
• Proven ability to lead technical discussions, mentor engineers, and drive best practices across delivery teams.
• Experience improving reliability through automated validations, data quality checks, and operational excellence practices.
Good to have skills:
Hadoop, HDFS, YARN, Airflow, HBase
Location: PAN INDIA
EXP:5-15 Years
Preferred candidate profile
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
