Experience
5 - 10 yrs
Salary (CTC)
₹5.4L - ₹6.9L
Job Location
Hyderabad, India
Vacancy
5
Designation
Pyspark Developer
Job Type
Not specified
Job Description
Job Locations : Bengaluru, Chennai, Hyderabad, Pune, Kolkata
Job Requirements:
- Experience with strong proficiency in Python and PySpark.
- Experience working on Spark SQL, RDD, and DataFrame APIs.
- Good understanding of Hadoop ecosystem (Hive, HDFS, YARN).
- Knowledge of data formats like Parquet, Avro, JSON, etc.
- Experience with SQL and writing efficient queries.
- Familiarity with job orchestration tools (Airflow, Oozie, or similar).
- Version control systems like Git.
- Exposure to cloud platforms (AWS/GCP/Azure) is a plus.
Key responsibilities:
- Design, develop, and maintain robust ETL/ELT pipelines using PySpark and other big data technologies. Optimize Spark jobs for performance and scalability.
- Implement data transformations, aggregations, and joins over large datasets. Perform batch and real-time data processing tasks.
- Collaborate with data scientists, analysts, and other engineers to understand requirements and deliver quality solutions. Integrate PySpark solutions with data warehouses (like Hive, Redshift, Snowflake) and other data stores.
- Write clean, maintainable, and well-documented code. Follow version control and CI/CD practices using tools like Git, Jenkins, or Azure DevOps.
- Troubleshoot data quality and performance issues. Monitor pipeline health and ensure SLAs are met.
- Work with tools and platforms like Hadoop, Hive, HDFS, AWS EMR, Databricks, or Azure Synapse as required.
