Job Description
We are looking for a highly skilled Senior Data Engineer with 7+ years of experience in developing and maintaining large-scale data processing solutions using Spark, PySpark, Scala, Python, Hive, and Apache Airflow. The ideal candidate should have strong expertise in distributed computing, ETL development, DataLake, workflow orchestration, and automation.
Hands-on experience with Cloudera CDP (On-Premise) environments is highly preferred. The candidate should also be comfortable using modern AI-assisted development tools such as Claude, Cursor, GitHub Copilot to improve development productivity and code quality.
We are looking for a highly skilled Big Data Engineer with expertise in Apache Spark, PySpark, Scala, Python, Hive, and Apache Airflow to build and support enterprise-scale data platforms.
The role requires hands-on experience with Hive SQL, partitioning, bucketing, performance tuning, and ETL workflows, along with strong knowledge of Spark SQL, DataFrames, Datasets, UDFs, and Spark optimization techniques. Experience with HDFS, Parquet, Cloudera CDP, and data modeling concepts such as Star and Snowflake schemas is essential.
Candidates should have experience with Git, GitLab, Maven/SBT, CI/CD pipelines.
Experience with GitHub Copilot, Cursor, or Claude, as well as exposure to Impala, Kafka, and Power BI, is a plus.
Key Responsibilities
- Design, develop, and maintain enterprise-scale Spark/PySpark data pipelines using Scala and Python.
- Work with HDFS, Parquet, and large-scale distributed data environments.
- Optimize Spark, Hive, and ETL workloads for performance, scalability, and reliability.
- Implement CI/CD pipelines and DevOps best practices using GitLab, Jenkins, Maven, and SBT.
- Utilize AI-powered development tools to improve code quality, engineering productivity, and development efficiency."
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
